Global foundation-model progress briefing
English Edition中文
Enter keywords to search ingested stories.

Today / Sunday, September 27, 2026

limbo logolimbo

Data updated

Aug 24, 06:22 PM

Live sources

17

Ingestion status

Live ingest

ResearchThe Decoder

Pew study confirms sharp rise of AI-written text on the web since ChatGPT's launch

Summary

The Pew Research Center analyzed nearly half a million English-language web pages for AI-generated content. More than a third of pages published since ChatGPT's launch show signs of machine-written text, and commercial .

Original Article

Captured source content or English translation, normalized into this reading format.

Read Source

Pew study confirms sharp rise of AI-written text on the web since ChatGPT's launch

Pew study confirms sharp rise of AI-written text on the web since ChatGPT's launch

  • In an analysis of nearly half a million English-language web pages, the Pew Research Center found that more than a third of pages published after ChatGPT's launch show signs of AI-generated text.
  • Texts from the Common Crawl web archive were analyzed using the Open Pangram detection tool. Commercial .com domains contain AI-generated text roughly ten times more often than .edu or .gov sites.
  • The analysis has limits, though, mainly because current detection tools can barely tell the difference between fully automated text and writing that was only partly AI-assisted.

The Pew Research Center analyzed nearly half a million English-language web pages for AI-generated content. Since ChatGPT launched, the share of machine-written text online has climbed sharply.

The texts came from theCommon Crawlweb archive and were checked for signs of machine authorship using the AI detection toolOpen Pangram. In a sample from July 2026, about 10 percent of all pages examined showed clear signs of AI authorship.

Filtering the sample to only include pages published after ChatGPT's release changes the picture dramatically. More than a third of those newer pages show signs of AI authorship,according to the analysis. The trend kicked off with ChatGPT in late 2022, and the share of likely AI-generated web content has climbed steadily ever since. Ad

The share of pages with AI-generated text has grown sharply since ChatGPT's release. | Image: Pew Research Center

About one in ten pages with a .com domain shows signs of AI authorship, while .org domains sit at 4.6 percent and .edu and .gov domains come in at only about 1 percent each. That makes commercial websites roughly ten times more likely to contain AI-written text than pages from schools or government agencies. Ad

Commercial .com sites are far more likely to use AI-generated text. | Image: Pew Research Cente

"Delve," em dashes, and Oxford commas are booming

Pew's analysis found several language patterns that have become much more common on the web since 2023. Em dashes now show up about twice as often as they did in 2023, and Oxford comma usage has jumped 63 percent.

Words, phrases, and punctuation patterns typical of AI text have spiked across the web. | Image: Pew Research Center

Certain AI-favorite words like "delve," "interplay," "testament," "pivotal," "landscape," "tapestry," "bolstered," "crucial," "meticulous," and "vibrant" have more than doubled in frequency. Negative parallelisms following the "it's not just X, it's Y" pattern have nearly tripled, though they remain rare in absolute numbers. A separate study looking at corporate PR documents found thatthis particular phrase quadrupledsince 2022. Ad

Astudy by Imperial College London, the Internet Archive, and Stanford Universityfrom April 2026 reached a similar conclusion, finding that roughly 35 percent of all newly published websites were fully or partly AI-generated. The researchers also found 33 percent higher semantic similarity between AI texts and a much more positive tone overall but cautioned that public perception of negative effects often goes well beyond what the data actually supports.

What counts as "AI text" remains fuzzy

There's a problem with this and similar studies, and with the public debate too. Nobody agrees on what "AI text" even means. The spectrum runs from fully automated content to human drafts polished with AI to texts where a model only stepped in for a few sentences. Ad

Open Pangram and similar models can, in my experience from testing hundreds of my own texts, only make a very rough call on whether a human or a machine likely wrote something. They can't reliably tell you how much AI was involved or at what stage, and they still misfire regularly. Yet these are very different ways of working. Ad

Public debate around AI text is growing more polarized, as the recentdiscussion about Anthropic's planned watermark for Claude outputmade clear. Using AI tools already carries a stigma that cuts both ways, somethingworkplace studies have documentedas well. Neither camp leaves much room for the messy reality of how people actually write with these tools. And since AI adoption isn't slowing down, figuring out that middle ground is going to matter more and more.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

Region

Europe

Heat Score

84

Category

Research

Language

en