DeepSeek V4.1 Flash achieved perfect results on an AI hacking benchmark, gaining code execution on all 11 vulnerable targets while keeping four fixed targets secure, completing the full attack run for $4.65 by leveraging cached tokens. The model demonstrated strong exploitation abilities across Grafana, Jenkins, and Nextcloud, finding both intended attack paths and alternative routes to achieve code execution.
GPT-5.6 Luna costs 28x less than GPT-6 Astra for code review ($0.0041 vs $0.113 per review) but finds fewer bugs: 69 verified bugs versus Astra's 92 across 50 pull requests, with a 24% false-positive rate compared to Astra's 4%. Luna is suitable for everyday correctness issues but unreliable for security-critical code like authentication and permissions.
Grafana released version 13.0.8, addressing three security vulnerabilities (CVE-2026-12704, CVE-2026-14199, CVE-2026-19475) and fixing a bug in the legacy version history page related to version dates and user display names.