Thứ Hai, 7 thg 9, 2026

Mô hình GPT-6 Astra vừa ghi nhận kết quả 14% trong bài benchmark MazeBench, cho thấy hiệu suất vượt trội gấp 7 lần so với đối thủ Claude Fable 5.1.

Nguồn chính
SOURCE@htihle15:21 7 thg 9Astra spent 60+ hours in this 3D open world spatial reasoning eval
SUPPORT@rohanpaul_ai22:30 7 thg 9reduced a projected 3B-token trajectory to about 350M tokens
SOURCE@kimmonismus14:35 7 thg 9Astra Pro scores 86.5%, almost tied with Claude Fable 5.1 at 86.6%
SUPPORT@reach_vb14:27 7 thg 9Astra is SoTA on MazeBench by a huge margin
SUPPORT@andrewcurran_14:47 7 thg 9massive jump with Astra