---
format: "aidr-story-markdown/v1"
id: "90e9195777129f14ef887f3cceff3f2366e703ff217860c1315ed9fd05078f78"
canonical_url: "https://aidr.today/90e91957?lang=en"
title: "Terminal Bench 4.0 Debuts 66 Tasks In Expanded AI Agent Evaluation"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-29T05:37:48.000Z"
category: "Research"
topics: ["agent","benchmark","open-source","reasoning"]
source_urls: ["https://huggingnews.com/ai/terminal-bench-40-debuts-66-tasks-in-expanded-ai-agent-evaluation-858bff43","https://x.com/ZixuanLi_/status/2093496352915562552","https://x.com/himanshustwts/status/2093510073129853241","https://x.com/ZixuanLi_/status/2093556798364008545","https://x.com/andykonwinski/status/2093549026298020157","https://x.com/jun_song/status/2093500903601147912"]
summary: "Developers released an updated version of the Terminal-Bench dataset and leaderboard to better align evaluation tools with current AI model development. The 4.0 release includes 66 tasks across science, software, ML, operations, security, hardware, and media, alongside a new Terminal-Bench-Science track for evaluating AI agents on research workflows. Early results show Opus 5 outperformed Fable 5, while GLM 5.3 emerged as the top open source model. Ryan Marten and Steven Dillmann pushed the update to recalibrate task resource assumptions for the leaderboard. The expansion comes as model post-training progress catches up with downstream use cases, reducing performance discrepancies for models from labs other than OpenAI or Anthropic."
---

# Terminal Bench 4\.0 Debuts 66 Tasks In Expanded AI Agent Evaluation

> [Open the canonical story](<https://aidr.today/90e91957?lang=en>)

**Published:** 2026-08-29T05:37:48.000Z
**Category:** Research
**Topics:** agent, benchmark, open\-source, reasoning

## Summary

Developers released an updated version of the Terminal\-Bench dataset and leaderboard to better align evaluation tools with current AI model development\. The 4\.0 release includes 66 tasks across science, software, ML, operations, security, hardware, and media, alongside a new Terminal\-Bench\-Science track for evaluating AI agents on research workflows\. Early results show Opus 5 outperformed Fable 5, while GLM 5\.3 emerged as the top open source model\. Ryan Marten and Steven Dillmann pushed the update to recalibrate task resource assumptions for the leaderboard\. The expansion comes as model post\-training progress catches up with downstream use cases, reducing performance discrepancies for models from labs other than OpenAI or Anthropic\.

## Sources

- [Story source](<https://huggingnews.com/ai/terminal-bench-40-debuts-66-tasks-in-expanded-ai-agent-evaluation-858bff43>)
- [Story source](<https://x.com/ZixuanLi_/status/2093496352915562552>)
- [Supporting source](<https://x.com/himanshustwts/status/2093510073129853241>)
- [Supporting source](<https://x.com/ZixuanLi_/status/2093556798364008545>)
- [Supporting source](<https://x.com/andykonwinski/status/2093549026298020157>)
- [Supporting source](<https://x.com/jun_song/status/2093500903601147912>)

