---
format: "aidr-story-markdown/v1"
id: "9e3ed547b6b0c99709e69d2e46b205c8bf64dba9349d557bfea728da815a9eb4"
canonical_url: "https://aidr.today/9e3ed547?lang=en"
title: "Training a model to identify AI-generated web content from structure alone"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-22T13:00:49.000Z"
category: "Research"
topics: ["detection","fine-tuning","benchmark","safety","open-source","llm","google","colab"]
source_urls: ["https://arxiv.org/abs/2609.15369","https://news.ycombinator.com/item?id=49800566","https://developers.googleblog.com/colab-is-now-part-of-your-google-ai-plan/"]
summary: "Word-level detectors identify unedited AI-generated text almost perfectly, but the literature documents their brittleness under rewording, and a word-level score neither characterizes a text nor identifies which AI model wrote it. We ask whether AI-generated text can be identified one level deeper, from structural signatures: how information is presented, in what order, with what evidence, and in what voice. We replicate StoryScope (Russell et al., 2026), which showed such patterns for AI-generated fiction, on commercial content: 2,250 pre-ChatGPT human blog posts from 268 company domains against 11,250 AI mirrors from five frontier models. A 214-feature instrument, applied by an LLM and validated in a human gold-annotation session (human-human kappa 0.928, human-model 0.946), detects AI posts from its 187 structural features alone at 98.0 macro-F1 on held-out companies, unchanged (98.1) when every AI post is reworded by its own model. The signal characterizes and attributes: AI posts share a tidy, self-announcing shape, 79.3% are attributed to the correct source against a 16.7% chance rate, and human posts occupy rare structural configurations. All effects replicate StoryScope's,…"
---

# Training a model to identify AI\-generated web content from structure alone

> [Open the canonical story](<https://aidr.today/9e3ed547?lang=en>)

**Published:** 2026-09-22T13:00:49.000Z
**Category:** Research
**Topics:** detection, fine\-tuning, benchmark, safety, open\-source, llm, google, colab

## Summary

Word\-level detectors identify unedited AI\-generated text almost perfectly, but the literature documents their brittleness under rewording, and a word\-level score neither characterizes a text nor identifies which AI model wrote it\. We ask whether AI\-generated text can be identified one level deeper, from structural signatures: how information is presented, in what order, with what evidence, and in what voice\. We replicate StoryScope \(Russell et al\., 2026\), which showed such patterns for AI\-generated fiction, on commercial content: 2,250 pre\-ChatGPT human blog posts from 268 company domains against 11,250 AI mirrors from five frontier models\. A 214\-feature instrument, applied by an LLM and validated in a human gold\-annotation session \(human\-human kappa 0\.928, human\-model 0\.946\), detects AI posts from its 187 structural features alone at 98\.0 macro\-F1 on held\-out companies, unchanged \(98\.1\) when every AI post is reworded by its own model\. The signal characterizes and attributes: AI posts share a tidy, self\-announcing shape, 79\.3% are attributed to the correct source against a 16\.7% chance rate, and human posts occupy rare structural configurations\. All effects replicate StoryScope's,…

## Sources

- [Story source](<https://arxiv.org/abs/2609.15369>)
- [Discussion](<https://news.ycombinator.com/item?id=49800566>)
- [Story source](<https://developers.googleblog.com/colab-is-now-part-of-your-google-ai-plan/>)

