---
format: "aidr-story-markdown/v1"
id: "d51bb6dd8c3c8fcf94ab2dae3c8c295d75fcc04d6a91fc63727aac6efd95e8aa"
canonical_url: "https://aidr.today/d51bb6dd?lang=en"
title: "OpenAI Releases GPT-6 Astra With 169 Epoch Record and Lower Monitorability"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-04T18:54:36.000Z"
category: "Models"
topics: ["openai","gpt","reasoning","benchmark","arc-agi-2","safety","agi","nvidia"]
source_urls: ["https://huggingnews.com/ai/openai-releases-gpt-6-astra-with-169-epoch-record-and-lower-monitorabili-af0159b4","https://x.com/Reuters/status/2095820415834927428","https://x.com/theinformation/status/2095874632565842027","https://x.com/kimmonismus/status/2095824384753549517","https://x.com/teortaxesTex/status/2095830049609879945","https://x.com/MadisonMills22/status/2095877628615954914","https://x.com/rasbt/status/2095871852711153742","https://x.com/AISecurityInst/status/2095885998949355689"]
summary: "The newest large language model from OpenAI provides an increase in interactive reasoning, reaching 98% on FrontierMath Tier 4. GPT-6 Astra scored 169 on the Epoch AI capability index, breaking the previous record of 163, and 84% on Mystery Game Puzzles compared to a previous high of 59%. On the ARC-AGI-3 benchmark, the model posted a 66% score, which rose to nearly 100% when paired with a continuous conversation harness at a cost of $360 per game. The model utilizes a reasoning technique that makes its internal processes harder for researchers to inspect, increasing the risk that the AI could evade human monitoring. This architectural trade-off aims to cut operational costs and improve coding output, though it hinders the detection of dangerous behavior. CEO Sam Altman stated this week that AI models are becoming \"superhuman\" in some areas, leaving the company to sail in \"unknown waters.\""
---

# OpenAI Releases GPT\-6 Astra With 169 Epoch Record and Lower Monitorability

> [Open the canonical story](<https://aidr.today/d51bb6dd?lang=en>)

**Published:** 2026-09-04T18:54:36.000Z
**Category:** Models
**Topics:** openai, gpt, reasoning, benchmark, arc\-agi\-2, safety, agi, nvidia

## Summary

The newest large language model from OpenAI provides an increase in interactive reasoning, reaching 98% on FrontierMath Tier 4\. GPT\-6 Astra scored 169 on the Epoch AI capability index, breaking the previous record of 163, and 84% on Mystery Game Puzzles compared to a previous high of 59%\. On the ARC\-AGI\-3 benchmark, the model posted a 66% score, which rose to nearly 100% when paired with a continuous conversation harness at a cost of $360 per game\. The model utilizes a reasoning technique that makes its internal processes harder for researchers to inspect, increasing the risk that the AI could evade human monitoring\. This architectural trade\-off aims to cut operational costs and improve coding output, though it hinders the detection of dangerous behavior\. CEO Sam Altman stated this week that AI models are becoming "superhuman" in some areas, leaving the company to sail in "unknown waters\."

## Sources

- [Story source](<https://huggingnews.com/ai/openai-releases-gpt-6-astra-with-169-epoch-record-and-lower-monitorabili-af0159b4>)
- [Story source](<https://x.com/Reuters/status/2095820415834927428>)
- [Supporting source](<https://x.com/theinformation/status/2095874632565842027>)
- [Supporting source](<https://x.com/kimmonismus/status/2095824384753549517>)
- [Supporting source](<https://x.com/teortaxesTex/status/2095830049609879945>)
- [Supporting source](<https://x.com/MadisonMills22/status/2095877628615954914>)
- [Supporting source](<https://x.com/rasbt/status/2095871852711153742>)
- [Story source](<https://x.com/AISecurityInst/status/2095885998949355689>)

