Executive Summary

On August 20, Pew Research Center published an analysis of nearly half a million webpages. The number the press carried was the headline one: more than a third of pages published since ChatGPT's release show signs of AI authorship. The part that touches day-to-day work sits further down the page, in the table that splits the same measure by domain.

In the six-month average at the January 2026 crawl point, 9.35% of .com pages showed signs of AI authorship. The share on .org was half that, 4.59%. On .edu it was 1.03%, and on .gov 0.76%. In the second half of 2022, when ChatGPT went public, all four domains were bunched near 1%. The total has grown over the three and a half years since, but what changed more is where that total lands.

For anyone training on crawl data, that stratification becomes a pipeline problem. A filter that sweeps a whole corpus behind one web-average figure runs without knowing where the contamination piled up. The per-domain source data Pew published alongside its chart puts numbers on why the filtering criterion has to move from the total to the source.

Pew Research Center logo — the organization behind this per-domain AI authorship study
▲ Pew Research Center, the organization behind this study | Source: Pew Research Center

Key Numbers

The first three figures give the size of the domain gap and the speed at which it opened. The last is the limit that needs the most care when this study gets quoted.

Sources: Pew Research Center chart source data and methodology (2026-08-20)

9.35%

of .com pages show signs of AI authorship

Six-month average at the January 2026 crawl point. In the same window .org was at 4.59%

0.76%

of .gov pages show the same signs

.edu is at 1.03%. The two domains contributed only 669 and 1,571 sampled pages

8.7x

growth on .com

From 1.07% in the July 2022 crawl, just before ChatGPT, to 9.35%. Over the same span .edu went from 0.64% to 1.03%

10-15%

of pages carry a detectable publication date

The widely quoted 35% comes from that subset. Pew wrote down that it should not be read as the web as a whole

1

The Four Domain Lines Came Apart in 2023

Pew drew a random sample of 10,000 English-language webpages from each of the 49 crawls Common Crawl produced between January 2021 and July 2026. The 490,000 pages went through a detection model, and the results were split by top-level domain and plotted as six-month averages. That time series is the most valuable part of the study. Plenty of outlets have already quoted the share at a single point in time, but the order in which those shares came apart exists only in the series.

Common Crawl logo — the open web archive this study's page sample was drawn from
▲ Common Crawl, the nonprofit that crawls and republishes the web every month | Source: Common Crawl

In the July 2022 crawl, the last one before ChatGPT went public, the readings were .com 1.07%, .org 0.63%, .edu 0.64% and .gov 0.25%. Four lines pressed against the floor, nearly on top of one another. From the first half of 2023, .com pulls away first, .org follows at roughly half the pace, and .edu and .gov drift around 1% for four straight years.

Signs of AI authorship by domain (six-month averages, %) 0 2 4 6 8 10 ChatGPT released Nov 2022 .com 9.35 .org 4.59 .edu 1.03 .gov 0.76 Jan 2021 Jan 2022 Jan 2023 Jan 2024 Jan 2025 Jan 2026 Common Crawl crawl date (Jan 2021 to Jan 2026)
▲ Plotted from the source data Pew Research Center published alongside its chart | Pebblous original graphic

Pew makes no claim about what caused the gap. What the study observes is correlation, and why only .edu and .gov stayed low falls outside its scope. One observable difference between the four does stand out, though. Addresses under .edu and .gov are not open to anyone who wants to register one, and documents posted there usually come attached to an affiliation, a named author and an institutional sign-off. The channels with more friction before publication are the ones where the accumulation stayed slow, and that much the data shows directly.

2

What Pew Actually Counted

The instrument was editlens_Llama-3.2-3B, an open-weight detection model released by Pangram. Feed it the body text of a page and it returns a score between 0 and 1, and Pew treated anything at 0.2 or above as a meaningful sign of AI authorship or editing. The threshold was not set at random. It was calibrated to line the open model up with the commercial one, according to the methodology. Across 62,370 pages from seven crawls, Open Pangram and the flagship commercial model Pangram 3.3 agreed in 96% of cases, with a Cohen's kappa of 0.61. Disagreements remained at the level of individual pages, but the aggregate shares the two models produced per crawl landed close together, and the shape of the upward trend was the same either way.

Pangram logo — the company behind the open-weight AI text detection model editlens_Llama-3.2-3B used in this study
▲ Pangram, maker of the open-weight detection model editlens_Llama-3.2-3B | Source: Pangram

The model never saw the full HTML of each page. Pew used the body text file that Common Crawl stores separately from the raw markup, the WET record that contains just the text of the page with the HTML code, images and other media stripped out. That is also the file most teams pick up first when they turn a web corpus into training data. The layer of text this study measured is the same one that reaches a training run.

The widely quoted 35% comes with a condition attached. Only 10% to 15% of pages in a given crawl sample carry a publication date field in their HTML. The 35% is the share within that subset whose date falls after ChatGPT's release. Pew did not bury this. The methodology says plainly that "the subset of pages with a detectable publication date is not a random subset of the web, so our post-ChatGPT estimate reflects the prevalence of AI authorship among dated content rather than the web as a whole." The caveat fell off in the headlines and the number traveled alone.

A caveat is not the same as a debunking. In the same document Pew notes that an independent study of newly published sites, built on Internet Archive data and a different methodology, arrived at the same 35%. Both figures describe documents that are verifiably new, and within that scope they support each other. The trouble starts when the value is carried over and quoted as a share of the web.

The same document holds one more caveat that matters even more for reading the domain gap. Open Pangram has a higher false-positive rate on pages created before AI use was common, and on the 2021 and 2022 crawls the model returned estimates of around 1%. Run the commercial model on those same pages and the values come out far lower. Do not read that as 1% of the pre-ChatGPT web having been written by AI. That value is the height of this model's false-positive floor.

The .edu reading of 1.03% and the .gov reading of 0.76% sit at exactly that floor, indistinguishable from it. The true share on those two domains could well be lower, which would make the tenfold gap against .com a lower bound. Sample size pushes in the same direction. In the January 2026 window the crawl captured 669 .gov pages and 1,571 .edu pages. A .gov share of 0.76% amounts to roughly five pages, and the line does swing hard, from 1.72% in July 2024 down to 0.76% in January 2026.

2.1The Web's Prose Moved Along With It

Separately from the detector scores, Pew counted the sampled text a second time, tallying how often certain marks and words appear per 10,000 words. Looking only at pages with publication dates after ChatGPT's release, and comparing the January 2023 point with the January 2026 point, all four measures climb.

Mark or vocabulary Jan 2023 Jan 2026 Change
Em dashes 5.79 11.19 1.9x
Oxford commas 34.04 55.51 1.6x
AI-typical vocabulary 11.94 26.02 2.2x
Negative parallelism 0.87 2.36 2.7x

Uses per 10,000 words, plotted as six-month averages. The vocabulary list is a fixed set of 27 items Pew defined in advance, including "delve," "interplay," "testament" and "bolstered." Negative parallelism means a comparison structured as "it's not just X, it's Y." Source: Pew Research Center chart source data

Pew did not count these marks only on the AI-written pages. It counted them across the whole sample, human-written pages included, so the average sentence on the web has drifted in that direction. The curves are not simple either. Em dash use actually sank into the low 4s through the second half of 2023 and into 2024 before turning steep in 2025 and doubling. Negative parallelism nearly tripled yet remains rare at 2.36 uses per 10,000 words, and Pew is explicit that no single one of these markers proves AI authorship on its own.

The style checklist Pebblous uses to review Korean drafts flags the same two habits, dash overuse and the "not simply X, but Y" construction. Seeing them confirmed statistically across the English-language web is useful, but the direction is what matters in practice. The more common these markers become, the less any judgment resting on them can carry, and the more weight shifts to the information sitting outside the sentence. Which address the document came from is one such signal.

3

Eight in Ten Sampled Pages Are .com

The source data Pew made downloadable next to the chart also carries per-domain sample counts. In the January 2026 window the four domains account for 49,719 pages, and 41,116 of them are .com. The rest is .org with 6,363, .edu with 1,571 and .gov with 669. More than eight in ten pages come from one domain. The table counts only pages that fell into those four top-level domains, so .net and country-code addresses are absent, but among the four, .edu and .gov together hold just 4.5%.

49,719 pages captured across the four domains, January 2026 window .com 41,116 pages (82.7%) .org 6,363 12.8% .edu 1,571 3.2% .gov 669 1.3%
▲ Sample-count columns (n_com, n_org, n_edu, n_gov) from Pew's chart source data, converted to shares | Pebblous original graphic

That composition sends you back to the two earlier numbers. Across the full 10,000-page sample from July 2026, 10% of pages showed signs of AI authorship, and .com on its own came in at 9.35%. The two are measured over different windows. The 10% is a single snapshot from one July 2026 crawl, and the 9.35% is a six-month average at the January 2026 point, so they are not values you can add or subtract. Still, their sitting side by side is no accident. The largest block in the sample is .com, which makes the figure everyone calls the web average close to a restatement of .com. However clean .edu and .gov may be, they carry no weight to pull the average down.

What the sample misses tilts the same way. To make it into Common Crawl, a page has to be publicly accessible. Pew's methodology notes that sites with paywalls or a login requirement are likely underrepresented in these samples. Subscription journalism, closed academic archives and documents behind institutional accounts were never counted. This study measured the open portion of the web, and any team building training data from the crawl inherits exactly the same portion.

Pretraining corpora derived from Common Crawl inherit that composition too. Contamination arrives shaped like a domain instead of settling over a corpus as an even fog, and how much of each source a corpus holds decides how contaminated that corpus actually is. Pebblous has written before about the problem of "clean data" being asserted without proof, and this dataset points to one minimum unit of that proof. Not the total, but the distribution across sources.

4

Filtering Works Only If You Stratify by Source First

That composition already spells out what happens when you design a filter around the single figure of a 10% web average. Sweep the whole corpus at one uniform strength on the assumption that contamination is spread evenly, and most of the .com documents where it actually collected survive, while institutional documents that had no problem to begin with get cut at the same rate. What remains is a corpus with roughly the same contamination rate as before, only smaller.

The point where practice can change sits earlier in the pipeline than most teams assume. If you keep the original URL alongside the text when you clean crawl data, the top-level domain is information you already hold. Stratifying by source before scoring each document, then applying different thresholds and different sampling rates per stratum, removes more contamination for the same budget. It also cuts the cost of running the detector, since there is no reason to spend the same compute on a stratum that holds almost nothing.

4.1Four Things to Check in Your Pipeline Today

Pew's study design leaves behind a few working rules for anyone handling a corpus.

  • Carry the source field through to the derived datasets. Once the URL or the top-level domain drops out at an intermediate stage, stratified filtering stops being possible from that moment on.
  • Keep detector scores as continuous values instead of freezing them into binary labels. Pew's 0.2 was itself chosen to calibrate against a commercial model, and moving the threshold changes the picture.
  • Record the sample count for each stratum. A share drawn from a stratum of a few hundred documents swings on a handful of cases, so those strata are safer with sampled review attached than with automated rules.
  • Measure your detector's false-positive floor first. Run documents from before AI use was common through the same model, treat the resulting value as that stratum's floor, and stop reading anything below it as signal.

That last rule applies anywhere a detector is used. The study Pebblous covered earlier, in which 21% of ICLR 2026 reviews showed signs of AI, used a detection model from the same vendor, and there too the confidence came from aggregate counts rather than from any individual document. A verdict on a single document is the output of a probabilistic tool, and the moment a pipeline hardens it into fact, a false positive becomes a label.

Editor's Note: An organization that manages data quality only in aggregate cannot see this kind of stratification. It can answer what percentage of the whole is a problem, but not where that problem concentrates, because nothing in the record says. That is why Pebblous asks about the manageability of sources alongside quality metrics when it talks about AI-Ready Data. This study shows why the question is practical, on the largest corpus there is. The same 10% is workable once you know where it settled, and once you do not, there is nothing left to do but shrink the total.

R

References

Research Reports

Tools and Data

  • 4.Pangram. "editlens_Llama-3.2-3B." Hugging Face (checked Aug 26, 2026). The open-weight AI detection model used in this study.
  • 5.Pangram. "Introducing Open Pangram." (checked Aug 26, 2026).
  • 6.Common Crawl. "Common Crawl." (checked Aug 26, 2026). The web archive the sample was drawn from.