---
format: "aidr-story-markdown/v1"
id: "12c7c4b93f1e3a449874069b9582785fa6a140290dd9946bd4dcebb519c16b4d"
canonical_url: "https://aidr.today/12c7c4b9?lang=en"
title: "AI Web Text Hits 31% and Harms Language Model Training"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-10-01T15:42:09.000Z"
category: "Research"
topics: ["llm","openai","agent","google"]
source_urls: ["https://marketbrief.now/ai/ai-web-text-hits-31percent-and-harms-language-model-training-7a3ac5f5","https://huggingnews.com/ai/ai-web-text-hits-31percent-and-harms-language-model-training-7a3ac5f5","https://x.com/TransluceAI/status/2105725928357937410","https://x.com/jackhcable/status/2105460442399437149","https://x.com/GerritD/status/2105467194276720731","https://huggingnews.com/cybersecurity/ai-agents-probed-white-house-and-department-of-war-websites-df053323","https://marketbrief.now/cybersecurity/ai-agents-probed-white-house-and-department-of-war-websites-df053323","https://huggingnews.com/ai/70percent-of-teens-use-ai-for-schoolwork-with-plunging-test-scores-027b8f87"]
summary: "A team of researchers analyzed the pretraining of 800 different models to determine how synthetic data affects the ability of AI to predict human text. By testing language models with parameter counts from 19.9M to 973M, the study found that incorporating AI-generated tokens into training sets can eventually increase loss on held-out human text, degrading overall performance. Pangram identifies 31.1% of FineWeb-filtered web tokens as AI-generated in August 2026, an increase from 27.5% in June 2026 and 10% in June 2024. For data-starved models, AI tokens initially improve accuracy before the benefit reverses into harm, while models with high budgets of human text experience immediate degradation. The researchers' new scaling law predicts these effects for models up to 3.6x larger with 41% lower error than the best previous estimates."
---

# AI Web Text Hits 31% and Harms Language Model Training

> [Open the canonical story](<https://aidr.today/12c7c4b9?lang=en>)

**Published:** 2026-10-01T15:42:09.000Z
**Category:** Research
**Topics:** llm, openai, agent, google

## Summary

A team of researchers analyzed the pretraining of 800 different models to determine how synthetic data affects the ability of AI to predict human text\. By testing language models with parameter counts from 19\.9M to 973M, the study found that incorporating AI\-generated tokens into training sets can eventually increase loss on held\-out human text, degrading overall performance\. Pangram identifies 31\.1% of FineWeb\-filtered web tokens as AI\-generated in August 2026, an increase from 27\.5% in June 2026 and 10% in June 2024\. For data\-starved models, AI tokens initially improve accuracy before the benefit reverses into harm, while models with high budgets of human text experience immediate degradation\. The researchers' new scaling law predicts these effects for models up to 3\.6x larger with 41% lower error than the best previous estimates\.

## Sources

- [Story source](<https://marketbrief.now/ai/ai-web-text-hits-31percent-and-harms-language-model-training-7a3ac5f5>)
- [Story source](<https://huggingnews.com/ai/ai-web-text-hits-31percent-and-harms-language-model-training-7a3ac5f5>)
- [Story source](<https://x.com/TransluceAI/status/2105725928357937410>)
- [Story source](<https://x.com/jackhcable/status/2105460442399437149>)
- [Supporting source](<https://x.com/GerritD/status/2105467194276720731>)
- [Story source](<https://huggingnews.com/cybersecurity/ai-agents-probed-white-house-and-department-of-war-websites-df053323>)
- [Story source](<https://marketbrief.now/cybersecurity/ai-agents-probed-white-house-and-department-of-war-websites-df053323>)
- [Story source](<https://huggingnews.com/ai/70percent-of-teens-use-ai-for-schoolwork-with-plunging-test-scores-027b8f87>)

