Pew Research published a study yesterday that puts a number on something we've all felt: 10% of all English-language webpages now show significant signs of AI authorship. Among pages published after ChatGPT launched, it's over one in three.
The methodology is solid. They pulled nearly half a million pages from Common Crawl spanning five years (2021–2026), ran them through Open Pangram — a detection model that flags linguistic patterns more common in AI text than human writing — and tracked the trend over time. The curve is steep and still curving up.
xychart-beta title "AI-Authored Webpages Over Time" x-axis ["Pre-ChatGPT", "2023", "2024", "2025", "Jul 2026"] y-axis "Percentage" 0 --> 35 line [1.2, 3.8, 6.1, 8.4, 10.0] line [2.0, 8.0, 15.0, 24.0, 34.0]
The gap between those two lines is the story. The total web is creeping up (10% feels low because old pages dominate archives). But the new web — the content being produced right now — is majority-automated in many categories. Filter by post-ChatGPT publication date and the number jumps to 34%. Filter by .com domains specifically and it's even higher.
.edu and .gov domains show far less AI content (4.6% and lower). Makes sense — those surfaces have editorial process, liability, and reputation at stake. .com has none of those constraints.
What this means
Search quality is about to get harder to maintain. Google's entire business model assumes the web is an honest signal of human interest. When a third of new content is optimized for engagement metrics by models that were trained on engagement-optimized content, the feedback loop tightens into a knot.
Detection is an arms race, not a solution. The same week this study lands, a tool called watermarks-remover went viral for stripping provenance signals from Claude, Gemini, and OpenAI outputs. You can't detect your way out of a problem where the countermeasure ships faster than the signal.
The real signal: We are past the point of being able to unsee AI content from the web. The internet's entropy has permanently increased. The question isn't how to filter it — it's how to build systems that work with higher noise floors. Search, recommendation, fact-checking, training data curation — all of these need to account for the fact that a significant and growing fraction of the corpus was written by something that didn't mean what it said.