Executive Summary
A tidy intuition spread quickly among content publishers: "If I block GPTBot and Google-Extended in robots.txt, I protect my data." The measured results point the other way. Large news publishers that blocked AI crawlers lost a meaningful share of their visitors, yet the rate at which they were cited inside AI answers barely dropped. Blocking looked like a defense, but the data suggests it is closer to self-harm — severing your own distribution channel. This report retraces why that paradox arises, and where content owners should actually be holding their leverage, using first-party numbers.
The clearest signal is the citation-retention rate after blocking. Depending on the bot, anywhere from 70.6% to 92.3% of sites that attempted to block were still cited in AI answers. The reason is structural. Blocking does not drain the reservoir of already-learned knowledge; it only shuts off the pipe that feeds new water in. Models build answers from frozen parameters and a separate search index, so changing robots.txt today leaves the standing water untouched. The size of the traffic loss is contested: an initial estimate of −23.1% for large publishers was revised down to roughly −7% on a weekly basis in the same authors' later draft.
So the destination of this report is not the normative prescription "don't block." It is a reallocation of leverage. If the core problem is the asymmetry — crawlers take the content but return almost no visits — then the lever that reverses it is not the block switch but an infrastructure where sources are tracked and settled. Instead of locking your data away, structure "where it came from and how it was used" so it becomes settleable. That is where bargaining power begins.
92.3%
Citation retention after blocking
Share of sites blocking Google-Extended still cited in AI answers. 70.6–92.3% by bot
~95%
Citations from blocking sites
Even sites blocking training bots made up ~95% of ChatGPT's citation sources
−23% → −7%
Large-publisher traffic loss
v1 monthly −23.1% → authors' later draft weekly ~−7%. Large publishers only
11,122 : 1
Anthropic crawl-to-referral
Over 11,000 pages crawled per visitor sent. Extreme vs. Google's 5:1
The Measured Paradox: Visitors Lost, Citations Kept
Start with the numbers on what blocking actually produced. When BuzzStream, a content-marketing tools company, analyzed 4 million AI citations and 3,600 prompts using the citation-tracking tool XOFU, it found that sites that blocked AI crawlers were mostly still cited afterward. Citation-retention rates ran ChatGPT-User 70.6%, OAI-SearchBot 82.4%, GPTBot 88.2%, and Google-Extended 92.3%. If the goal of blocking was "get my content out of AI answers," that goal was largely unmet.
The more counterintuitive part is the citation share held by blocking sites. Even the sites that blocked training bots still accounted for about 95% of ChatGPT's citation sources. There are concrete cases too. CNBC.com, which blocked several bots at once, still logged 1,298 citations inside the analyzed dataset, and Yahoo.com, which blocked Google-Extended, retained roughly 30,000 citations. The door was locked, yet the AI kept reaching in and pulling out what was inside.
The visitor side tells a different story, and here is the number this report handles most carefully. An initial draft (v1, 2025-12-31) by Junhao Zhao of Rutgers Business School and Ron Berman of Wharton, analyzing how publishers responded to generative AI, reported that large publishers that blocked crawlers lost 23.1% of total traffic on a monthly basis (human traffic −13.9%). But when the same authors re-measured with an identical methodology on a weekly basis in the v4 revision (2026-04, SSRN), the figure fell to roughly −7%. Widen the sample to the top 500 and mid-sized publishers actually gained traffic, while some publishers saw losses clear enough that they rolled their blocking rules back.
So "23%" should be read not as an absolute but as an early estimate limited to large publishers. Whether the headline number is 23.1% or 7%, this report's argument does not hinge on the value. The direction is the same in both: visitors are lost while citations remain. The magnitude of the loss is contested; the fact that loss and retention diverge is not.
The chart below puts that asymmetry on a single screen. On the left is the traffic large publishers lost (both versions shown); on the right is the citation-retention rate that survived blocking, by bot. The gap between the bar lengths is where this report begins.
Figure 1. The asymmetry of blocking — visitor loss vs. citation retention. Source: original chart reconstructed by Pebblous from Zhao & Berman (Rutgers/Wharton, v1 2025-12 / v4 2026-04) for traffic and BuzzStream × Citation Labs (2026-03, 4M citations, 3,600 prompts) for citation retention.
One caveat is worth stating. BuzzStream's operational definition of "citation" is not airtight: it counted the links and sources listed in AI responses, without controlling for how or where they were displayed. Even so, the direction held consistently across all four bots, and two different samples — the traffic paper and the citation survey — point to the same conclusion. On that basis, the proposition that "citations survive blocking" itself stands firm.
Why Blocking Can't Stop Citations: Reservoir and Pipe
The previous section showed what happened; now we ask why. The answer, it turns out, fits into a single metaphor. Blocking does not empty the reservoir of training data; it only closes the pipe that feeds new water in.
There are two paths by which a citation in an AI answer gets made. One is knowledge frozen into the parameters of an already-trained model; the other is a separate search index referenced at answer time. What robots.txt controls is only the crawl stage that collects data. It cannot intervene at the moment an answer is generated. So even if you block crawlers today, the water already standing in the reservoir — snapshots learned in the past and copies held in the Common Crawl archive — does not retroactively drain out. And this reservoir is not shallow. By Mozilla Foundation's analysis, roughly 64% of major LLMs use a filtered version of Common Crawl in pre-training. Content archived once on the web becomes a water source shared not by a single model but by the entire ecosystem, so even if your site shuts its door now, those copies are already scattered across the parameters of many models.
The existence of this "already-standing water" shows up in the data, too. In the BuzzStream survey, about 15% of the content AI cited was published before ChatGPT even existed. Block the crawlers or not, the model keeps drawing up water it drank long ago. The diagram below is that structure — the block switch closes one pipe, but the reservoir level stays put.
Figure 2. Reservoir-and-pipe concept. Source: original diagram built by Pebblous from Longpre et al., "Consent in Crisis" (Data Provenance Initiative, arXiv:2407.14933) and BuzzStream's (2026) analysis of citation timing.
Here is one piece of academic grounding. The Data Provenance Initiative's "Consent in Crisis" study (Longpre et al.) tracked how fast robots.txt restrictions were surging across the web. Among the domains most important for training, more than 25% of tokens were newly restricted within a single year. But that surge carries a lag and an asymmetry. Restrictions apply only to data that will be collected going forward; they are not retroactive to already-distributed archives (for example, the widely circulated Common Crawl snapshots in WARC format). In other words, "block from now on" and "remove what was already learned" are entirely different problems. The former is possible with robots.txt; the latter is, for all practical purposes, not.
Let us be clear about one thing. This section is not trying to detail the internal search-and-citation pipeline of any specific model. The confirmed scope reaches only the observation that "citation is not contingent on real-time crawling." Any finer mechanism — which model references which index at which moment — exceeds the data this report has in hand, so we rest only on the observed results here.
The Hidden Bias of Blocking: Who Locks the Door
So far the question has been whether blocking helps or hurts an individual site; here we widen the lens a step. If blocking does not happen evenly, that unevenness itself distorts how representative of the web the model learns from — because the side that locks its door is skewed.
A paper by Paul Bouchaud and Pedro Ramaciotti of France's ISC-PIF and médialab (arXiv:2510.09031), analyzing crawler-restriction patterns across the top one million websites since 2023, reports three kinds of gap. First, by scale: about 25% of the top 1,000 sites block, but widen to the full top million and it falls to about 10% — bigger sites have more capacity and more legal resources to block. Second, by outlet type: news outlets overall block GPTBot at 34.2%, but among outlets with high fact-checking credibility that rises to about 55%. Third — and this is the paper's central finding — the gap by political leaning.
Center-leaning outlets blocked at about 58%, while right-leaning outlets blocked at only about 4.1%. A gap of more than tenfold. The chart below lines up all three gaps; all three axes point the same way — the higher-quality, more centrist, and larger a site is, the more it blocks.
Figure 3. Three block-rate gaps. Source: reconstructed by Pebblous from Bouchaud & Ramaciotti, arXiv:2510.09031 (2025-10). Scale, outlet, and political-leaning axes all from the same paper.
The authors' interpretation runs as follows. Centrist, high-fact-check outlets, ahead on resources and legal awareness, block more aggressively. As a result, the share of centrist sources in the training corpus falls relatively, while the share of hyperpartisan sources rises relatively. Politically skewed vocabulary was in fact observed to be over-sampled, and the authors close with a warning that "active intervention in training-data curation is needed."
This direction is confirmed independently, too. A paper by a different team using a similar method (arXiv:2510.10315, "Is Misinformation More Open?") reports that 60.0% of high-credibility sites block at least one AI crawler, whereas only 9.1% of misinformation sites do. High-credibility sites disallow an average of 15.5 AI user-agents; misinformation sites average fewer than one. Swap the leaning axis for a credibility axis and the conclusion is the same. The side locking its door is, by and large, the side producing verifiable, neutral content.
Here the problem with blocking moves beyond one site's traffic profit or loss. When the places making good data drop out more often, the average quality of the web the model trains on and cites from goes down. Blocking looks like a personal choice to "protect what's mine," but in aggregate it produces a collective outcome: it erodes the representativeness of the training corpus.
The Legal and Standards Gap in the Block Switch: A Convention No One Has to Obey
If blocking can't stop citations and instead biases the corpus, then we should ask how trustworthy the robots.txt tool is in the first place. To put the conclusion first: this switch sits on softer ground than people assume. It has no legal force and no enforceability as a standard.
robots.txt is only a voluntary protocol that has run since 1994; there is no penalty for ignoring it. In fact, bots' robots.txt non-compliance rate was 12.9% in Q1 2025, nearly quadrupling from 3.3% in the prior period. Declare "I've blocked it" all you like — if the other side ignores it, that's that. The Data Provenance Initiative's "Consent in Crisis" study adds another layer. Many cases were found where a site's robots.txt clauses and its Terms of Service (ToS) contradict each other — one permits while the other forbids. With the channel for expressing intent split in two, a crawler has no reliable way to read "what this site actually wants."
Even so, site operators attempt fairly sophisticated decisions: "keep my search visibility but block only AI training." In the same Consent in Crisis data, the conditional probability that a site also blocked each organization's crawler — given that it blocked any crawler at all — makes this separated intent stand out sharply.
| Crawler organization | Prob. of co-blocking | Type |
|---|---|---|
| OpenAI (GPTBot) | 91.5% | AI training / collection |
| Common Crawl | 83.4% | Public archive |
| Anthropic | 83.4% | AI training / collection |
| Google-Extended | 72.0% | AI training (separate from search) |
| Cohere | 52.3% | AI training / collection |
| Meta | 52.2% | AI training / collection |
| Internet Archive | 32.3% | Non-profit archive |
| Google Search | 17.1% | Pure search indexing |
Table 1. Conditional probability of being "co-blocked," by organization. Source: Longpre et al., "Consent in Crisis" (Data Provenance Initiative, arXiv:2407.14933, Table 3). The probability that each organization was also blocked, given that a site blocked any crawler at all.
The gap between OpenAI at the top (91.5%) and Google Search at the bottom (17.1%) is the heart of this table. Operators block AI-training crawlers aggressively while leaving Google Search, which handles search indexing, mostly open. It is evidence that they are deciding with a clear distinction between "AI learning my writing to replace it with a summary" and "search bringing people to my writing." The problem is that no means of enforcement follows this careful distinction. Google separating training and search into distinct tokens is a step forward, but compliance with even that option is left to the operator's goodwill.
So publishers began moving beyond robots.txt. In the first half of 2026, the U.S. News/Media Alliance (representing some 900 publishers) sent a formal letter to Common Crawl, and Digital Content Next (AP, NYT, NBCUniversal, Bloomberg, NPR, and others) went public with a cease-and-desist. The legal argument they advanced compresses into a single sentence: "copyright is not an opt-out regime." To use the work you must get permission in advance — not take it first and be asked to remove it afterward. On that logic, the opt-out model of robots.txt is itself at odds with publishers' rights claims.
Just how slow and incomplete after-the-fact removal is, the European cases show. In November 2025 the Dutch copyright body BREIN had more than 2 million paywalled items removed from training datasets, and one Danish removal request took about six months from intake to response, after which only roughly half of the targeted material was actually erased. Pulling specific content out of an already-distributed archive is closer to putting spilled water back in the cup.
The block switch has no enforceability, no retroactivity, and no clear legal footing. It is an expression of intent, not a contract. That is precisely why settlement infrastructure has emerged: an attempt to move from a take-it-or-leave-it protocol to a structure where ignoring it either blocks access or triggers a charge. Only that settlement, too, is still just half open.
The Real Leverage: Not Blocking, but Settlement
If blocking is not a defense, what is the lever content owners should actually be holding? To answer that, look again at the essence of the problem. What publishers are really losing is not "being trained on" but the asymmetry of value taken while traffic is not returned. And that asymmetry is measured cleanly by a single metric: the crawl-to-referral ratio.
By Cloudflare's measurements on its own network (through which roughly 20% of the entire web passes), how many times a crawler scrapes pages for every visitor it sends back splits dramatically by operator. Traditional search — Google — sends one visitor per five crawls (5:1), but OpenAI runs 1,700:1 and Anthropic 11,122:1. That means it scrapes over 11,000 pages and returns a single visitor. The chart below puts the gap on a log scale; without a log axis the Google bar would be too small to see against the others.
Figure 4. Crawl-to-referral ratio comparison (log scale). Source: reconstructed by Pebblous from Cloudflare Radar (Matthew Prince, mid-2026 data). The gap widened from the July 2025 reading; figures are unified to the latest data so that values from different points in time are not placed side by side as absolutes.
This asymmetry is becoming the default rather than the exception. In mid-2026, automated requests made up 57.5% of HTML traffic on Cloudflare's network — the first time machines surpassed humans. Yet the share of those AI crawl requests that led to an actual human visit was just 2.6%. The reader of the web is increasingly a machine, and that machine sends almost no people back to the source. That blocking can't stop this flow, the previous three sections already showed. The remaining option, then, is "make them settle for what they take."
Settlement is only half open
Settlement infrastructure has already entered the commercial stage. Cloudflare's pay-per-crawl lets a site charge at least $0.01 per page when a crawler takes it, via an HTTP 402 (Payment Required) and an Ed25519-signed request. Time, Condé Nast, The Atlantic, AP, Reddit, and Stack Overflow are among the early participating publishers. It replaces the "block or open" binary with a third option: "price it and open."
But this settlement applies only at the collection (crawl) stage. You can put a price on the moment a crawler takes a page, but there is still no infrastructure to price the value at the inference stage — the value generated every time an already-trained model produces an answer. You can meter the water as it fills the reservoir, but the settlement each time it comes out of the tap is empty. And as we saw, a substantial share of citations comes precisely from that "already-filled reservoir." The real battleground for settlement is not collection but inference.
So what content owners should be demanding is not tighter blocking rules but the preconditions that make settlement possible. First, structured license metadata that can track what was used, when, and by whom. Second, a settlement protocol that distinguishes collection from inference and prices each stage. Third, lineage and provenance transparency that can verify where data came from and how it flowed. Only with these three in place can an asymmetry like "11,122:1" be put on the negotiating table as a number.
The conclusion is not a norm but a reallocation of strategy. Not "don't block," but: put down the weak lever of blocking and pick up the strong levers of settlement and provenance. Lock your data away and you sever only the distribution channel; the reservoir stays full. Structure your data to be settleable and, conversely, a cost attaches to the taker and bargaining power accrues to the owner. Turning the hand that locks the door into the hand that sets the price — that is where this report arrives, by way of the data.
Editor's Note: Why We Watch This Topic
There is a reason the team that wrote this report has watched this topic for so long. The place Pebblous works is exactly where this article's conclusion points. The prescription of the final section — "not blocking but settlement, and provenance as the precondition for settlement" — is the same problem we grapple with every day.
First, we want to revisit the finding of Section 3. The story that the average quality of the training corpus drops when the places making good data lock their doors more often is a data-quality problem beyond any one site's traffic gain or loss. When what gets learned into a model becomes uncontrollably biased, that bias hardens into the model's internal representations. These papers, in effect, prove with data why managing the provenance and representativeness of training data matters.
And the three preconditions Section 5 called for — structured license metadata, a stage-by-stage settlement protocol, and lineage/provenance transparency — all converge on one state: "data whose origin and use are traceable." That is precisely the state Pebblous calls 'AI-Ready Data,' and the quality scoring and lineage management DataClinic performs are the tools for reaching it. To reverse an asymmetry like 11,122:1, "what was used, when, and by whom" has to be recorded first — and that record is provenance.
This is not a piece written to sell that tool. We simply wanted to convey, with data, that while those who hold data believe a single line of robots.txt has protected them, the real leverage is being built somewhere else. The way past the paradox where blocking becomes a loss lies not in locking data away but in making it settleable. We intend to keep watching the next phase of this topic: how inference-stage settlement infrastructure gets built.
Related Pebblous reports on this topic: Wikipedia and the Draining of the AI Commons, RSL Content Licensing, The Cost of Auditing Open-Dataset Licenses, AI Distillation and Data Sovereignty.
References
Academic papers
- 1.Bouchaud, P., & Ramaciotti, P. (2025). "Web Crawler Restrictions, AI Training Datasets & Political Biases." arXiv:2510.09031. (ISC-PIF / médialab) — primary source for the three block-rate gaps by scale, outlet, and political leaning.
- 2.(2025). "Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web." arXiv:2510.10315. — high-trust 60.0% vs. misinformation 9.1% block rate; cross-validation on the credibility axis.
- 3.Longpre, S. et al. (2024). "Consent in Crisis: The Rapid Decline of the AI Data Commons." arXiv:2407.14933. (Data Provenance Initiative) — surging token restrictions, robots.txt/ToS mismatch, per-organization co-blocking probability (Table 3).
- 4.Zhao, J. (Rutgers), & Berman, R. (Wharton). (2025–2026). "Strategic Response of News Publishers to Generative AI." SSRN. — large-publisher traffic loss, v1 −23.1% (monthly) → v4 ~−7% (weekly). Note the figures differ by version.
Policy & statistics
- 5.Nero, V. (BuzzStream) × Citation Labs. (2026). "Do News Sites That Block AI Bots Still Get Cited?" — 4M citations, 3,600 prompts (XOFU); citation-retention rate 70.6–92.3% by bot.
- 6.Reuters Institute. (2026). "Journalism, Media, and Technology Trends and Predictions 2026" (survey of 280 media executives). Reuters Institute. — background trend of AI-search-driven traffic decline (distinct from blocking effects).
- 7.Mozilla Foundation. (2024). "Common Crawl and the AI training data ecosystem" (64% of LLMs use Common Crawl). Mozilla Foundation.
- 8.Prince, M. (Cloudflare). (2025–2026). "Introducing pay-per-crawl" & Cloudflare Radar. — crawl-to-referral ratio (Google 5:1 / OpenAI 1,700:1 / Anthropic 11,122:1), pay-per-crawl ($0.01/page, HTTP 402, Ed25519).
Industry coverage
- 9.ppc.land. "Blocking AI crawlers doesn't stop citations — new data shows why." — entry point for this report.
- 10.digitalapplied.com. "Publishers vs Common Crawl: AI Training Data Showdown 2026." — the 2026 Common Crawl negotiation phase and the cease-and-desist narrative.