OpenAI's newest artificial intelligence model achieved a score of 1,797 on the Code Arena WebDev benchmark, taking the top spot from Anthropic's Claude Fable 5.1. The GPT-6 Astra (Max) model matches competitive pricing at $40 per million tokens and beats OpenAI's previous GPT-5.6 Sol entry by 180 points. In PINNACLE enterprise testing, the model completed 279 of 280 multi-step jobs at a cost of $1.51 per successful task, which is 39% less than Claude's $2.46. At its maximum reasoning setting, Astra fabricated no answers on questions that could not be answered from provided information.
The update moves AI capability toward direct computer control of browsers and terminals to execute multi-step workflows, increasing token volume and compute demand per job. Initial access is limited to cybersecurity customers, with API and AWS integrations following. This agentic shift puts structural pressure on seat-based software pricing by automating workflows that previously required individual human user seats, shifting value toward compute infrastructure and data governance platforms.
Key sources
- SOURCE@arena“best-performing model at $40/Mtoken”x.com
- SUPPORT@arena“agentic coding workflows that require multi-step reasoning and tool use”x.com
- SUPPORT@wallstengine“Completed 279 of 280 jobs from start to finish”x.com
- SUPPORT@saxena_puru“moving from copilot assistance (generating text and code) towards direct computer control”x.com
- SUPPORT@spicey_lemonade“massive 90% expected performance”x.com
- SOURCE@j_dekoninck“on average, it uses only 12k tokens per question for BrokenArXiv and ArXivMath”x.com
- SUPPORT@firstadopter“astra is *incredible* for anything that requires vision or computer use. no other model comes close”x.com
- SUPPORT@jukan05“That’s a staggering amount of compute.”x.com