Epoch AI's Benchmark Reviews initiative evaluates the quality of tools used to measure artificial intelligence capabilities. The first assessment of 15 benchmarks designated 9 as flawed, while 4 were verified and 2 lacked enough information for a review. Findings showed that 45.5% of tasks in Terminal Bench 4.0 were broken, and 46% of randomly sampled questions in the HLE benchmark were also found to be faulty.

Audit results for DeepSWE 1.1 identified a bug capable of breaking grading for every task and 23 false negatives out of 131 tasks, which creates an artificial performance ceiling of 79.6%. The verifier in that benchmark discards agent changes to some test files without notifying the model, leading to irrelevant failures. Epoch AI publishes full assessments for all verified benchmarks to create a consistent quality standard and incentivize higher industry benchmarks.

Sign in to suggest edits

Key sources

  1. SOURCE@epochairesearch“launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information”x.com
  2. SUPPORT@epochairesearch“any errors that exist do not substantially affect the results”x.com
  3. SUPPORT@tamaybes“In Terminal Bench 4.0 we found 45.5% of tasks to be broken”x.com
  4. SUPPORT@nrehiew_“on DeepSWE, Epoch found 23 false negatives out of 131 tasks”x.com
  5. SUPPORT@charliermarsh“in DeepSWE v1.1, the verifier discards the agent's changes to some test files”x.com
  6. SUPPORT@emollick“tells us how bad the state of benchmarking is, and how terrible some of our favorite benchmarks are”x.com
  7. SUPPORT@epochairesearch“Benchmarks assess AI capabilities, but the benchmarks themselves vary substantially in quality”x.com
  8. SUPPORT@epochairesearch“To avoid conflicts of interest we do not review Epoch-created benchmarks”x.com
Markdown