← Back to live feed1 story
Developers released an updated version of the Terminal-Bench dataset and leaderboard to better align evaluation tools with current AI model development. The 4.0 release includes 66 tasks across science, software, ML, operations, security, hardware, and media, alongside a new Terminal-Bench-Science track for evaluating AI agents on research workflows. Early results show Opus 5 outperformed Fable 5, while GLM 5.3 emerged as the top open source model.
Ryan Marten and Steven Dillmann pushed the update to recalibrate task resource assumptions for the leaderboard. The expansion comes as model post-training progress catches up with downstream use cases, reducing performance discrepancies for models from labs other than OpenAI or Anthropic.
Sign in to suggest edits
Key sources
- SOURCE@zixuanli_“Benchmark iteration is catching up with model development”x.com
- SUPPORT@himanshustwts“GLM 5.3 is best open source model out there sharing very close boundary with Fable 5”x.com
- SUPPORT@zixuanli_“pushed a version update to the Terminal-Bench dataset and leaderboard”x.com
- SUPPORT@andykonwinski“benchmark for evaluating AI agents on research workflows across scientific do…”x.com
- SUPPORT@jun_song“Those benchmark scores are from before the nerfs, and right now both are heavily degraded”x.com