← Back to live feed1 story
A benchmark for evaluating AI agents on research workflows across scientific domains
Sign in to suggest edits
Key sources
- DISCUSSIONmatt_dnews.ycombinator.com
A benchmark for evaluating AI agents on research workflows across scientific domains