---
format: "aidr-story-markdown/v1"
id: "ffe46162c5139c10bf7725f6f60b3266e4e7799dde2dc26227d752fe896649b5"
canonical_url: "https://aidr.today/ffe46162?lang=en"
title: "Grok 4.7 Tops First Enterprise AI Cyber Defense Index"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-28T16:15:12.000Z"
category: "Models"
topics: ["xai","agent"]
source_urls: ["https://marketbrief.now/ai/grok-47-tops-first-enterprise-ai-cyber-defense-index-e741529f","https://huggingnews.com/ai/grok-47-tops-first-enterprise-ai-cyber-defense-index-e741529f"]
summary: "Artificial Analysis partnered with Nvidia and IBM to release a new set of benchmarks measuring the ability of AI agents to find and patch software vulnerabilities. MiMo-V2.6-Pro and Grok 4.7 (xhigh) lead the results with scores of 56, while GPT-6 Luna and GLM-5.3-Flash follow with 53 and 50, respectively. Several frontier models, including Claude Opus 5.5 and Gemini 3.8 Flash, trail the leaders by 19 to 31 points after declining between 32% and 38% of the tasks on safety grounds. The evaluation uses three frameworks: CWE-Bench-AA for language-specific patching, DeepsecBench-AA for vulnerability discovery, and CyberGym-E2E-AA for memory safety fixes. In the CyberGym test, models such as GPT-6 Sol and GPT-6 Astra refused every task, while Claude Fable 5.1 declined 99% of assignments. The Cyber Index Alliance, which includes Vercel and CollinearAI, aims to establish an industry standard for defense as AI models gain more advanced cyber offense capabilities."
---

# Grok 4\.7 Tops First Enterprise AI Cyber Defense Index

> [Open the canonical story](<https://aidr.today/ffe46162?lang=en>)

**Published:** 2026-09-28T16:15:12.000Z
**Category:** Models
**Topics:** xai, agent

## Summary

Artificial Analysis partnered with Nvidia and IBM to release a new set of benchmarks measuring the ability of AI agents to find and patch software vulnerabilities\. MiMo\-V2\.6\-Pro and Grok 4\.7 \(xhigh\) lead the results with scores of 56, while GPT\-6 Luna and GLM\-5\.3\-Flash follow with 53 and 50, respectively\. Several frontier models, including Claude Opus 5\.5 and Gemini 3\.8 Flash, trail the leaders by 19 to 31 points after declining between 32% and 38% of the tasks on safety grounds\. The evaluation uses three frameworks: CWE\-Bench\-AA for language\-specific patching, DeepsecBench\-AA for vulnerability discovery, and CyberGym\-E2E\-AA for memory safety fixes\. In the CyberGym test, models such as GPT\-6 Sol and GPT\-6 Astra refused every task, while Claude Fable 5\.1 declined 99% of assignments\. The Cyber Index Alliance, which includes Vercel and CollinearAI, aims to establish an industry standard for defense as AI models gain more advanced cyber offense capabilities\.

## Sources

- [Story source](<https://marketbrief.now/ai/grok-47-tops-first-enterprise-ai-cyber-defense-index-e741529f>)
- [Story source](<https://huggingnews.com/ai/grok-47-tops-first-enterprise-ai-cyber-defense-index-e741529f>)

