Thursday, Oct 1, 2026

A team of researchers analyzed the pretraining of 800 different models to determine how synthetic data affects the ability of AI to predict human text. By testing language models with parameter counts from 19.9M to 973M, the study found that incorporating AI-generated tokens into training sets can eventually increase loss on held-out human text, degrading overall performance.

Pangram identifies 31.1% of FineWeb-filtered web tokens as AI-generated in August 2026, an increase from 27.5% in June 2026 and 10% in June 2024. For data-starved models, AI tokens initially improve accuracy before the benefit reverses into harm, while models with high budgets of human text experience immediate degradation. The researchers' new scaling law predicts these effects for models up to 3.6x larger with 41% lower error than the best previous estimates.

Sign in to suggest edits

Key sources

  1. SOURCEmarketbrief.now
  2. SOURCEhuggingnewshuggingnews.com
  3. SOURCE@transluceai“AI agents used aggressive, non-hacking tactics against government websites, including the White House, the Department of War, and several U.S. states”x.com
  4. SOURCE@jackhcable“disclosing new evidence of AI agents probing and attempting rudimentary vulnerability exploits against U.S. and Canadian government agencies”x.com
  5. SUPPORT@gerritd“Add Canada to the list of governments that have seen attempted hacks by agents”x.com
  6. SOURCEhuggingnewshuggingnews.com
  7. SOURCEmarketbrief.now
  8. SOURCEhuggingnewshuggingnews.com
Markdown