Method
What is measured
How much English-language news text published on the open web shows the fingerprints of AI writing, week by week, compared with news from before ChatGPT (before November 2022). The result is a trend for the whole population of sampled articles, with a 95% confidence range.
Where the documents come from
Common Crawl CC-NEWS, a public crawl of news sites. Each week we draw random pieces of its files, extract the article text, keep English documents with high language-detection confidence, and cap how many come from one site. The raw record location is stored with every document. The older weeks (back to 2019) are a smaller, separate “retrospective” sample, used only as the baseline and never mixed with the weekly data.
The two numbers
Word-shift index. Studies of scientific writing found words that became much more common after ChatGPT, such as “delve”, “underscore” and “intricate” (Kobak et al., Science Advances 2025; Liang et al., ICML 2024). We count a fixed list of them and divide by the rate in the pre-2022-11 baseline. 1.00 means no change; the range comes from resampling the documents.
Leftover-phrase floor. The share of documents containing text a person pasted without reading, such as “As an AI language model” or “Regenerate response”. It is a minimum. Most AI text has no such trace, so the true share is higher by an unknown amount.
What is not measured
- No per-document verdicts. Detectors that label one article as “AI” or “human” make errors that fall on real writers, especially non-native English writers. Estimates for a whole population avoid that, so we publish none for single articles or sites.
- The word list comes from science writing. Whether it shifts the same way in news is part of what is being measured, and models change their habits over time. A rising index means more of the shift, not a count of AI articles.
- Other languages, and sites that are not news, are out of scope for this version.
- Text that was edited by a person, or written by a model with a different style, is mostly invisible to both numbers.
Versions
Methodology version in use: v1.1. Changes to sampling, language handling, text extraction, word lists or phrases each get a new version, and old documents can be re-analysed under a new one. Limits: text extraction is a simple heuristic, and CC-NEWS covers news-type sites, not the whole web.