---
format: "aidr-story-markdown/v1"
id: "dd4fdc0d0a117d52e65e95cc6d5b40b39d206c90744ca92c80852ddc8c57d7c5"
canonical_url: "https://aidr.today/dd4fdc0d?lang=en"
title: "OpenAI GPT-6 Astra Beats Human Reasoning Baseline on SimpleBench with 86.5% Score"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-07T16:44:39.000Z"
category: "Models"
topics: ["openai","gpt","reasoning","benchmark"]
source_urls: ["https://huggingnews.com/ai/openai-gpt-6-astra-beats-human-reasoning-baseline-on-simplebench-with-86-587225f8","https://x.com/kimmonismus/status/2096970502111662376","https://x.com/reach_vb/status/2096968605518692367","https://x.com/AndrewCurran_/status/2096973604890247296","https://x.com/nicdunz/status/2097101602649739751","https://x.com/nicdunz/status/2097102538088542410","https://x.com/scaling01/status/2097032985019023606","https://x.com/scaling01/status/2097139066370207868"]
summary: "OpenAI released a new set of benchmark results for its GPT-6 Astra model, highlighting its ability to solve spatial logic and social scenarios. Astra Pro posted a score of 86.5% on SimpleBench, slightly behind Claude Fable 5.1's 86.6% but above the human baseline of 83.7%. The model also secured the state-of-the-art position on MazeBench, a 3D open world spatial reasoning evaluation. SimpleBench targets common logic problems that have repeatedly caused failures in previously released large language models. Current test results show a gap between the 86.5% AI score and the 83.7% human average. Some early evaluators described the performance as a \"massive jump\" in capability, though they observed that AI advancement remains jagged across different task types."
---

# OpenAI GPT\-6 Astra Beats Human Reasoning Baseline on SimpleBench with 86\.5% Score

> [Open the canonical story](<https://aidr.today/dd4fdc0d?lang=en>)

**Published:** 2026-09-07T16:44:39.000Z
**Category:** Models
**Topics:** openai, gpt, reasoning, benchmark

## Summary

OpenAI released a new set of benchmark results for its GPT\-6 Astra model, highlighting its ability to solve spatial logic and social scenarios\. Astra Pro posted a score of 86\.5% on SimpleBench, slightly behind Claude Fable 5\.1's 86\.6% but above the human baseline of 83\.7%\. The model also secured the state\-of\-the\-art position on MazeBench, a 3D open world spatial reasoning evaluation\. SimpleBench targets common logic problems that have repeatedly caused failures in previously released large language models\. Current test results show a gap between the 86\.5% AI score and the 83\.7% human average\. Some early evaluators described the performance as a "massive jump" in capability, though they observed that AI advancement remains jagged across different task types\.

## Sources

- [Story source](<https://huggingnews.com/ai/openai-gpt-6-astra-beats-human-reasoning-baseline-on-simplebench-with-86-587225f8>)
- [Story source](<https://x.com/kimmonismus/status/2096970502111662376>)
- [Supporting source](<https://x.com/reach_vb/status/2096968605518692367>)
- [Supporting source](<https://x.com/AndrewCurran_/status/2096973604890247296>)
- [Story source](<https://x.com/nicdunz/status/2097101602649739751>)
- [Supporting source](<https://x.com/nicdunz/status/2097102538088542410>)
- [Supporting source](<https://x.com/scaling01/status/2097032985019023606>)
- [Supporting source](<https://x.com/scaling01/status/2097139066370207868>)

