A Stanford-led community effort released Terminal-Bench-Science v0.1 to evaluate AI agents on research workflows across scientific domains. The initial benchmark contains 70 tasks, and Claude Opus 5 solved approximately 30% of them.

The team behind the original Terminal-Bench collaborated with scientific experts from research institutions worldwide and received sponsorship from Bespoke Labs. The initiative aims to develop reliable AI partners that allow scientists to prioritize human-centric work.

Sign in to suggest edits

Key sources

  1. SOURCE@modal“Terminal-Bench set the standard for evaluating long-horizon agents”x.com
  2. SUPPORT@madiator“it will spur progress in science!”x.com
Markdown