OpenAI released a new set of benchmark results for its GPT-6 Astra model, highlighting its ability to solve spatial logic and social scenarios. Astra Pro posted a score of 86.5% on SimpleBench, slightly behind Claude Fable 5.1's 86.6% but above the human baseline of 83.7%. The model also secured the state-of-the-art position on MazeBench, a 3D open world spatial reasoning evaluation.
SimpleBench targets common logic problems that have repeatedly caused failures in previously released large language models. Current test results show a gap between the 86.5% AI score and the 83.7% human average. Some early evaluators described the performance as a "massive jump" in capability, though they observed that AI advancement remains jagged across different task types.
Key sources
- SOURCE@kimmonismus“Astra Pro scores 86.5%, almost tied with Claude Fable 5.1 at 86.6%”x.com
- SUPPORT@reach_vb“Astra is SoTA on MazeBench by a huge margin”x.com
- SUPPORT@andrewcurran_“massive jump with Astra”x.com
- SOURCE@nicdunz“first model we've tested that seems to be able to handle basically any prompt”x.com
- SUPPORT@nicdunz“solving puzzles that require juggling 5 moving parts with ease”x.com
- SUPPORT@scaling01“Astra spent 60+ hours in this 3D open world spatial reasoning eval”x.com
- SUPPORT@scaling01“okay that is a disgusting mog if I've ever seen one”x.com
- SOURCE@htihle“Astra spent 60+ hours in this 3D open world spatial reasoning eval”x.com