LukiWebCo Dead Internet Tracker

Method

What is measured

How much English-language news text published on the open web shows the fingerprints of AI writing, week by week, compared with news from before ChatGPT (before November 2022). The result is a trend for the whole population of sampled articles, with a 95% confidence range.

Where the documents come from

Common Crawl CC-NEWS, a public crawl of news sites. Each week we draw random pieces of its files, extract the article text, keep English documents with high language-detection confidence, and cap how many come from one site. The raw record location is stored with every document. The older weeks (back to 2019) are a smaller, separate “retrospective” sample, used only as the baseline and never mixed with the weekly data.

The two numbers

Word-shift index. Studies of scientific writing found words that became much more common after ChatGPT, such as “delve”, “underscore” and “intricate” (Kobak et al., Science Advances 2025; Liang et al., ICML 2024). We count a fixed list of them and divide by the rate in the pre-2022-11 baseline. 1.00 means no change; the range comes from resampling the documents.

Leftover-phrase floor. The share of documents containing text a person pasted without reading, such as “As an AI language model” or “Regenerate response”. It is a minimum. Most AI text has no such trace, so the true share is higher by an unknown amount.

What is not measured

Versions

Methodology version in use: v1.1. Changes to sampling, language handling, text extraction, word lists or phrases each get a new version, and old documents can be re-analysed under a new one. Limits: text extraction is a simple heuristic, and CC-NEWS covers news-type sites, not the whole web.