OpenAI's newest model completed a 60 hour evaluation of long horizon spatial reasoning without relying on external Python code to solve the environment. GPT-6 Astra recorded a 14% result on the MazeBench benchmark, locating 100 gems across 200 rooms in a process that often requires more than 100 moves. This performance is 7 times the 2% score achieved by Claude Fable 5.1 on the same no-code track.
The model's result outperformed the 13% score of GPT-5.6 Sol's code enabled run, demonstrating an ability to maintain complex plans internally. To optimize the processing, Astra planned 10 to 20 moves ahead in batched actions, compressing a projected 3 billion token trajectory to 350 million tokens. Persistent failures on genuinely 3D puzzles suggest the gain is primarily in planning rather than complete spatial understanding.
Key sources
- SOURCE@htihle“Astra spent 60+ hours in this 3D open world spatial reasoning eval”x.com
- SUPPORT@rohanpaul_ai“reduced a projected 3B-token trajectory to about 350M tokens”x.com
- SOURCE@kimmonismus“Astra Pro scores 86.5%, almost tied with Claude Fable 5.1 at 86.6%”x.com
- SUPPORT@reach_vb“Astra is SoTA on MazeBench by a huge margin”x.com
- SUPPORT@andrewcurran_“massive jump with Astra”x.com