Original briefings. Zero spin.
Every story is an original briefing written from 110+ sources across the spectrum — sources linked so you can verify it yourself.
Study: More Than a Third of New Web Pages Now Show Signs of AI Authorship

The English-language internet is being handed over to machines, and the shift happened fast.
A peer-reviewed study led by researchers Jonas Dolezal, Sawood Alam, Mark Graham, and Maty Bohacek, from Imperial College London, Stanford University, and the Internet Archive, found that roughly 35% of websites published by mid-2025 qualify as AI-generated or AI-assisted. More than 20% of those pages showed no meaningful human editing at all, according to the study. Before ChatGPT's public launch in November 2022, that number was effectively zero.
The team built its sample from the Wayback Machine, pulling a stratified set of archived websites from mid-2022 through mid-2025. They ran the pages through Pangram v3, a detector trained to catch linguistic fingerprints left behind by large language models.
Pew Research, using the same Pangram detection technology, reached a nearly identical number through a separate methodology. Pew pulled almost half a million English-language pages from the Common Crawl archive covering the past five years, then isolated a 10,000-page sample from July 2026. When Pew filtered out pages published before ChatGPT existed, signs of AI authorship showed up in 35% of what remained, according to the report released Thursday.
Pew's data adds a detail the academic study didn't dig into: which corners of the internet are most affected. Pages ending in .com showed AI authorship at roughly 10 times the rate of .edu or .gov domains, both of which sat around 1%, Pew found. Nonprofit .org domains landed at 4.6%. Government and university sites, in other words, are still mostly human-written. The commercial web is where the machines have moved in.
What the data does and doesn't show
The Imperial College/Stanford/Internet Archive researchers specifically tested four fears people commonly raise about AI content: more factual errors, fewer outbound links, less stylistic variety, and lower semantic density in longer pieces. None of the four showed a statistically significant correlation with AI prevalence, according to the study.
This finding cuts against the instinct to assume AI content is automatically worse or riskier. The researchers did find two things that were statistically significant: AI-generated pages showed 33% higher semantic similarity to each other compared to human-written pages, and positive sentiment scores ran 107% higher on AI content. Translation: the AI web is more repetitive and relentlessly upbeat, even if it's not necessarily less accurate or less linked.
The researchers also surveyed 853 U.S. adults about what they believe AI content is doing to the internet. A majority believed in all six negative outcomes they were asked about, including the four the quantitative data couldn't back up. People assume AI content is riddled with errors and thin on substance. The study's own numbers say that assumption isn't holding up, at least not yet and not on the specific dimensions they measured.
Concerns about AI content may still be valid. A reasonable skeptic would point out that "no statistically significant correlation" on four narrow metrics doesn't prove AI content is trustworthy or well-sourced, just that it doesn't obviously fail on those particular tests. The study didn't claim to settle whether AI-written pages are accurate, only that accuracy didn't measurably decline as AI prevalence rose in this sample. That's a narrower claim than "AI content is fine," and it shouldn't be stretched further than the data supports.
The bot-on-bot problem
TechCrunch's coverage flagged a detail that deserves more attention than it got: this study measures what's being published, not who's reading it. Cloudflare has reported that bot web traffic recently overtook human web traffic, arriving at that milestone faster than the company had projected. Put the two data points together and a chunk of the internet is now bots writing content that other bots are reading, with no human in the loop on either end.
Neither study makes a claim about intent or malice here. Nobody's accusing publishers of some coordinated scheme. Content mills, SEO farms, and legitimate outlets alike are using AI tools because they're fast and cheap, and search engines haven't fully adapted their ranking systems to sort quality AI content from filler.
The researchers' proposed fixes are narrow and technical rather than alarmist: cryptographic verification of human authorship, and search or platform algorithms that deliberately weight toward diverse, human-generated content. Neither fix exists at scale yet. Google, Bing, and the major AI chatbot makers have not announced firm timelines for deploying either approach, based on the available reporting. Until that changes, the trendline in this study points one direction: up.
Pangram's detection method is not infallible. Both Pew and the academic researchers acknowledge that AI-detection tools can misclassify human-written text as machine-generated, and vice versa. At the scale involved here, though, both research teams argue the numbers are directionally reliable even if any single classification isn't perfect.
Sources used for this briefing
This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.