Mô hình GPT-6 Astra vừa ghi nhận kết quả 14% trong bài benchmark MazeBench, cho thấy hiệu suất vượt trội gấp 7 lần so với đối thủ Claude Fable 5.1.

Đăng nhập để góp ý, chỉnh sửa

Nguồn chính

  1. SOURCE@htihle“Astra spent 60+ hours in this 3D open world spatial reasoning eval”x.com
  2. SUPPORT@rohanpaul_ai“reduced a projected 3B-token trajectory to about 350M tokens”x.com
  3. SOURCE@kimmonismus“Astra Pro scores 86.5%, almost tied with Claude Fable 5.1 at 86.6%”x.com
  4. SUPPORT@reach_vb“Astra is SoTA on MazeBench by a huge margin”x.com
  5. SUPPORT@andrewcurran_“massive jump with Astra”x.com
Bản Markdown