Anthropic released new performance data for Claude Fable 5.1 showing it outperformed human participants on the SimpleBench benchmark, the first time an AI has exceeded that baseline. The model's Max version reached 1,765 points in the Code Arena: WebDev to rank first overall, jumping from 8th place in the previous Fable 5 generation with a 137 point increase.

Fable 5.1 holds the top rank in five of six categories, including Gaming and Simulations, while placing 8th in Data & Analytics. The system is available for direct testing in the Arena interface through Sept. 5 at 8 a.m. PT, after which it remains accessible only in Battle and Agent modes.

Sign in to suggest edits

Key sources

  1. SOURCE@scaling01“outscores humans on SimpleBench”x.com
  2. SUPPORT@lentils80“first ever model to beat the human baseline score in SimpleBench”x.com
  3. SOURCE@arena“massive point gains in Gaming (+234 pts), Simulations (+196 pts), and Reference-Based Design (+160 pts)”x.com
  4. SUPPORT@arena“available in Direct Mode through Sept. 5 at 8am PT”x.com
Markdown