OpenAI released a new set of benchmark results for its GPT-6 Astra model, highlighting its ability to solve spatial logic and social scenarios. Astra Pro posted a score of 86.5% on SimpleBench, slightly behind Claude Fable 5.1's 86.6% but above the human baseline of 83.7%. The model also secured the state-of-the-art position on MazeBench, a 3D open world spatial reasoning evaluation.

SimpleBench targets common logic problems that have repeatedly caused failures in previously released large language models. Current test results show a gap between the 86.5% AI score and the 83.7% human average. Some early evaluators described the performance as a "massive jump" in capability, though they observed that AI advancement remains jagged across different task types.

Sign in to suggest edits

Key sources

  1. SOURCE@kimmonismus“Astra Pro scores 86.5%, almost tied with Claude Fable 5.1 at 86.6%”x.com
  2. SUPPORT@reach_vb“Astra is SoTA on MazeBench by a huge margin”x.com
  3. SUPPORT@andrewcurran_“massive jump with Astra”x.com
  4. SOURCE@nicdunz“first model we've tested that seems to be able to handle basically any prompt”x.com
  5. SUPPORT@nicdunz“solving puzzles that require juggling 5 moving parts with ease”x.com
  6. SUPPORT@scaling01“Astra spent 60+ hours in this 3D open world spatial reasoning eval”x.com
  7. SUPPORT@scaling01“okay that is a disgusting mog if I've ever seen one”x.com
  8. SOURCE@htihle“Astra spent 60+ hours in this 3D open world spatial reasoning eval”x.com
Markdown