---
format: "aidr-story-markdown/v1"
id: "f6580b2cdbf30e14d15a0bd1bbcbe32e0d73ca1c3e6f45943ed2fede6e5cf3d2"
canonical_url: "https://aidr.today/f6580b2c?lang=en"
title: "GPT-6 Astra Scores 14% on MazeBench, 7x Lead Over Claude Fable 5.1"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-07T23:35:26.000Z"
category: "Models"
topics: ["gpt-6","maze-bench","spatial-reasoning","openai","competitive-benchmark","claude"]
source_urls: ["https://huggingnews.com/ai/gpt-6-astra-scores-14percent-on-mazebench-7x-lead-over-claude-fable-51-cfc2c8a9","https://x.com/htihle/status/2096982137412931850","https://x.com/rohanpaul_ai/status/2097090235104735624","https://x.com/kimmonismus/status/2096970502111662376","https://x.com/reach_vb/status/2096968605518692367","https://x.com/AndrewCurran_/status/2096973604890247296"]
summary: "OpenAI's newest model completed a 60 hour evaluation of long horizon spatial reasoning without relying on external Python code to solve the environment. GPT-6 Astra recorded a 14% result on the MazeBench benchmark, locating 100 gems across 200 rooms in a process that often requires more than 100 moves. This performance is 7 times the 2% score achieved by Claude Fable 5.1 on the same no-code track. The model's result outperformed the 13% score of GPT-5.6 Sol's code enabled run, demonstrating an ability to maintain complex plans internally. To optimize the processing, Astra planned 10 to 20 moves ahead in batched actions, compressing a projected 3 billion token trajectory to 350 million tokens. Persistent failures on genuinely 3D puzzles suggest the gain is primarily in planning rather than complete spatial understanding."
---

# GPT\-6 Astra Scores 14% on MazeBench, 7x Lead Over Claude Fable 5\.1

> [Open the canonical story](<https://aidr.today/f6580b2c?lang=en>)

**Published:** 2026-09-07T23:35:26.000Z
**Category:** Models
**Topics:** gpt\-6, maze\-bench, spatial\-reasoning, openai, competitive\-benchmark, claude

## Summary

OpenAI's newest model completed a 60 hour evaluation of long horizon spatial reasoning without relying on external Python code to solve the environment\. GPT\-6 Astra recorded a 14% result on the MazeBench benchmark, locating 100 gems across 200 rooms in a process that often requires more than 100 moves\. This performance is 7 times the 2% score achieved by Claude Fable 5\.1 on the same no\-code track\. The model's result outperformed the 13% score of GPT\-5\.6 Sol's code enabled run, demonstrating an ability to maintain complex plans internally\. To optimize the processing, Astra planned 10 to 20 moves ahead in batched actions, compressing a projected 3 billion token trajectory to 350 million tokens\. Persistent failures on genuinely 3D puzzles suggest the gain is primarily in planning rather than complete spatial understanding\.

## Sources

- [Story source](<https://huggingnews.com/ai/gpt-6-astra-scores-14percent-on-mazebench-7x-lead-over-claude-fable-51-cfc2c8a9>)
- [Story source](<https://x.com/htihle/status/2096982137412931850>)
- [Supporting source](<https://x.com/rohanpaul_ai/status/2097090235104735624>)
- [Story source](<https://x.com/kimmonismus/status/2096970502111662376>)
- [Supporting source](<https://x.com/reach_vb/status/2096968605518692367>)
- [Supporting source](<https://x.com/AndrewCurran_/status/2096973604890247296>)

