← Back to live feed1 story
A Stanford-led community effort released Terminal-Bench-Science v0.1 to evaluate AI agents on research workflows across scientific domains. The initial benchmark contains 70 tasks, and Claude Opus 5 solved approximately 30% of them.
The team behind the original Terminal-Bench collaborated with scientific experts from research institutions worldwide and received sponsorship from Bespoke Labs. The initiative aims to develop reliable AI partners that allow scientists to prioritize human-centric work.
Sign in to suggest edits