Artificial Analysis partnered with Nvidia and IBM to release a new set of benchmarks measuring the ability of AI agents to find and patch software vulnerabilities. MiMo-V2.6-Pro and Grok 4.7 (xhigh) lead the results with scores of 56, while GPT-6 Luna and GLM-5.3-Flash follow with 53 and 50, respectively. Several frontier models, including Claude Opus 5.5 and Gemini 3.8 Flash, trail the leaders by 19 to 31 points after declining between 32% and 38% of the tasks on safety grounds. The evaluation uses three frameworks: CWE-Bench-AA for language-specific patching, DeepsecBench-AA for vulnerability discovery, and CyberGym-E2E-AA for memory safety fixes. In the CyberGym test, models such as GPT-6 Sol and GPT-6 Astra refused every task, while Claude Fable 5.1 declined 99% of assignments. The Cyber Index Alliance, which includes Vercel and CollinearAI, aims to establish an industry standard for defense as AI models gain more advanced cyber offense capabilities.

Sign in to suggest edits

Key sources

  1. SOURCEmarketbrief.now
  2. SOURCEhuggingnewshuggingnews.com
Markdown