---
format: "aidr-story-markdown/v1"
id: "2edb5e5ea73a3165d52a9b958464850d8fd18650a27c7a5c7ecd1cf70ed668a0"
canonical_url: "https://aidr.today/2edb5e5e?lang=en"
title: "Artificial Analysis Zeroes AI Agent Scores in First Reward Hacking Correction"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-26T02:47:59.000Z"
category: "Research"
topics: ["agent","benchmark","safety"]
source_urls: ["https://huggingnews.com/ai/artificial-analysis-zeroes-ai-agent-scores-in-first-reward-hacking-corre-e9c579fe"]
summary: "Terminal-Bench v2.1 enables AI agents to browse the public web, a feature that has led some models to retrieve benchmark solutions from their training data. To…"
---

# Artificial Analysis Zeroes AI Agent Scores in First Reward Hacking Correction

> [Open the canonical story](<https://aidr.today/2edb5e5e?lang=en>)

**Published:** 2026-08-26T02:47:59.000Z
**Category:** Research
**Topics:** agent, benchmark, safety

## Summary

Terminal\-Bench v2\.1 enables AI agents to browse the public web, a feature that has led some models to retrieve benchmark solutions from their training data\. To…

## Sources

- [Story source](<https://huggingnews.com/ai/artificial-analysis-zeroes-ai-agent-scores-in-first-reward-hacking-corre-e9c579fe>)

