---
format: "aidr-story-markdown/v1"
id: "e2c32579ca3943549b495ba7cfccbe66e39dcfee2221b05ff75b5fe93ac63f42"
canonical_url: "https://aidr.today/e2c32579?lang=en"
title: "Claude Opus 5 Solves Only 30% of New Scientific AI Benchmark"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-27T20:55:33.000Z"
category: "Research"
topics: ["anthropic","claude","benchmark","agent"]
source_urls: ["https://huggingnews.com/ai/claude-opus-5-solves-only-30percent-of-new-scientific-ai-benchmark-d3363d27","https://x.com/modal/status/2093046013640450225","https://x.com/madiator/status/2093054302634062216"]
summary: "A Stanford-led community effort released Terminal-Bench-Science v0.1 to evaluate AI agents on research workflows across scientific domains. The initial benchmark contains 70 tasks, and Claude Opus 5 solved approximately 30% of them. The team behind the original Terminal-Bench collaborated with scientific experts from research institutions worldwide and received sponsorship from Bespoke Labs. The initiative aims to develop reliable AI partners that allow scientists to prioritize human-centric work."
---

# Claude Opus 5 Solves Only 30% of New Scientific AI Benchmark

> [Open the canonical story](<https://aidr.today/e2c32579?lang=en>)

**Published:** 2026-08-27T20:55:33.000Z
**Category:** Research
**Topics:** anthropic, claude, benchmark, agent

## Summary

A Stanford\-led community effort released Terminal\-Bench\-Science v0\.1 to evaluate AI agents on research workflows across scientific domains\. The initial benchmark contains 70 tasks, and Claude Opus 5 solved approximately 30% of them\. The team behind the original Terminal\-Bench collaborated with scientific experts from research institutions worldwide and received sponsorship from Bespoke Labs\. The initiative aims to develop reliable AI partners that allow scientists to prioritize human\-centric work\.

## Sources

- [Story source](<https://huggingnews.com/ai/claude-opus-5-solves-only-30percent-of-new-scientific-ai-benchmark-d3363d27>)
- [Story source](<https://x.com/modal/status/2093046013640450225>)
- [Supporting source](<https://x.com/madiator/status/2093054302634062216>)

