OpenAI's newest model completed a 60 hour evaluation of long horizon spatial reasoning without relying on external Python code to solve the environment. GPT-6 Astra recorded a 14% result on the MazeBench benchmark, locating 100 gems across 200 rooms in a process that often requires more than 100 moves. This performance is 7 times the 2% score achieved by Claude Fable 5.1 on the same no-code track.

The model's result outperformed the 13% score of GPT-5.6 Sol's code enabled run, demonstrating an ability to maintain complex plans internally. To optimize the processing, Astra planned 10 to 20 moves ahead in batched actions, compressing a projected 3 billion token trajectory to 350 million tokens. Persistent failures on genuinely 3D puzzles suggest the gain is primarily in planning rather than complete spatial understanding.

Sign in to suggest edits

Key sources

  1. SOURCE@htihle“Astra spent 60+ hours in this 3D open world spatial reasoning eval”x.com
  2. SUPPORT@rohanpaul_ai“reduced a projected 3B-token trajectory to about 350M tokens”x.com
  3. SOURCE@kimmonismus“Astra Pro scores 86.5%, almost tied with Claude Fable 5.1 at 86.6%”x.com
  4. SUPPORT@reach_vb“Astra is SoTA on MazeBench by a huge margin”x.com
  5. SUPPORT@andrewcurran_“massive jump with Astra”x.com
Markdown