Executive Summary
On August 18, arXiv published a field experiment that put two opposite versions of Google Search in front of real users: one with the AI answers stripped out, and one where AI answers were all you got. Researchers at the University of Pennsylvania and Northeastern University randomly assigned 1,100 people in the United States to three conditions and observed them for ten days. This is not a vendor announcement and not a traffic tally. It is a preregistered causal estimate.
In the condition that hid AI Overviews and AI Mode, the rate of clicks leaving search for other sites rose by 8.8 percentage points. In the condition that routed every search to AI Mode, it fell by 18.8 points. That much was expected. What was not expected is that in the same AI Mode condition, trust in the information Google provided, satisfaction, and perceived usefulness all fell along with it.
The common assumption is that users gain roughly what content suppliers lose. This experiment did not find that trade. This article reads the result not as a search marketing story but as a question about the supply of the source material that AI systems learn from and cite.
Key numbers
The first two numbers are what each condition did to clicks reaching publishers. The next two show which way users themselves moved under enforced AI Mode. All four are measured against the control group, which saw ordinary Google Search.
Source: Wang et al. (2026), arXiv:2608.18352
+8.8pp
Click-through, AI answers hidden
Clicks leaving search for other sites, LATE estimate
−18.8pp
Click-through, AI Mode enforced
Same measure, p<0.001
−0.34 pts
Trust in information on Google
7-point scale, AI Mode condition
33.6%
Negative impressions of AI Mode
29.3% positive, 309 open-ended responses
Splitting the search box three ways
How much search traffic AI summaries have absorbed has mostly been argued from publishers' own referral dashboards. Those dashboards show a decline, but they cannot isolate a cause. Ranking algorithms changed over the same months, seasonality moved, and reading habits moved with them. This experiment tried to separate the causal part by manipulating the search page itself inside each participant's browser.
Google does not accept the dashboard reading. As the paper summarizes the company's position, AI features are complements rather than substitutes, overall referral traffic has held steady, click quality has improved, and longer and more complex queries have opened up new ways for sites to be surfaced. When publisher counts and a platform's rebuttal describe the same phenomenon differently, random assignment is one of the few instruments that can say which account holds.
Participants installed a Chrome extension and were randomly assigned to three conditions. The first hid AI Overviews and AI Mode. The second changed nothing and served as the control. The third routed every search to AI Mode. Three days of baseline were followed by seven days of treatment, ten days in all. Of the 1,444 people who enrolled, 1,100 searched at least once during the experiment period and 956 finished the post-experiment survey.
Recruitment ran in waves between March 17 and March 19, 2026, through the research platform Prolific and a Northeastern work-study program. Only US residents who used Chrome and Google Search as their primary tools were admitted. Seven hypotheses and the analysis plan were preregistered on OSF before the experiment began, and the aggregated data and analysis code are posted on Harvard Dataverse. Attrition did not differ significantly across the three conditions.
The diagram below shows the three conditions together with the change in click-through each one produced. The middle line is the control group, and the distance above or below it is the change relative to that control.
The researchers also measured whether each manipulation reached the screen. In the AI Mode condition, 94.7% of searches were successfully routed to AI Mode, so that arm worked as designed. The condition that hid AI answers did not. Google rolled out a change to the AI Overviews HTML during the study, which broke the extension's logic for finding and hiding them. A hiding success rate of about 90% on the first day fell to roughly zero by the end, averaging 51.1% overall. Effects for that arm are therefore reported as local average treatment effects, estimated on the share of AI Overviews actually hidden. The preregistration had specified this fallback in advance.
Hide the answer and the clicks come back
With AI Overviews hidden, click-through rose by 8.8 percentage points, with a 95% confidence interval of 2.3 to 15.3 points and p = 0.008. Click-through here is the share of searches that produced a click out to a site beyond Google. Read it as the share of searches where the user traveled to the site that supplied the underlying material instead of stopping at the answer on the results page.
One condition has to travel with that number. Because the hiding logic worked only about half the time, the estimate computed on assignment alone is 2.7 percentage points and is not statistically significant. The 8.8-point figure scales the effect to the AI Overviews that were genuinely hidden. The authors flag in their limitations that this assumption can break. The direction, though, is clear. The more the AI summary disappeared from the page, the more often people left for the web.
The news part deserves a careful reading. The share of users clicking through to news sites did rise in this condition, but the effect did not survive counting clicks per day or classifying news domains a different way. That is why the authors call it partial support. What can be stated firmly is about the web at large rather than news outlets in particular. Take the AI answer away and the total volume of traffic search passes outward goes up.
Force AI Mode and the clicks fall across the board
The manipulation in the other direction was far larger. With every search routed to AI Mode, click-through fell by 18.8 percentage points. The decline was not concentrated in one sector. The share of users clicking through to news sites fell 12.5 points, Reddit 21.2 points, and Wikipedia 9.9 points. Ad clicks fell 42.7 points, because AI Mode did not surface ads at the time of the experiment.
The direction of the session metrics is the more striking part. The researchers had hypothesized that a conversational interface would raise the number of searches, and instead daily search sessions fell by 0.92. Minutes per session rose by 0.43 in exchange. People started searching less often, stayed longer once they did, and left less.
How large that drop is depends on the baseline. During the three days before treatment, participants averaged 4.0 search sessions, 9.2 searches, and 4.1 clicks per day. A loss of 0.92 sessions is close to a quarter of the sessions they used to run. Among heavy users, those with more than four sessions a day, the drop widened to 2.01 sessions, and this was the only moderator effect that remained significant after correcting for multiple comparisons.
The result was not comfortable for Google either. The share of participants in the AI Mode condition who used Bing, DuckDuckGo, or Yahoo rose by 11.2 points, and their stated intention to switch to Bing if AI Mode were permanent rose 1.25 points on a 7-point scale. A design that holds the answer on the page did not hold on to the people reading it.
Users lost something too
The debate about AI in search has mostly been framed as a trade-off. Users save time, publishers lose traffic, and the remaining question is how to compensate the loss. This experiment tested the first half of that framing, and the answer did not match the assumption.
Search research has a name for the assumption. When a user finds the answer on the results page and leaves without clicking anything, that state is called good abandonment. No click occurred, but the user got what they came for, so it is not read as a bad signal. The paper sets three interests side by side on top of that concept. User convenience, publisher sustainability, and Google's consolidation of informational authority do not easily align.
In the AI Mode condition, trust in the information Google provided fell 0.34 points on a 7-point scale. Satisfaction fell 0.73 standard deviations, perceived usefulness 0.59, and the sense of agency over one's own search 0.66. Perceived personalization and relevance fell 0.42 standard deviations, an outcome the researchers had predicted would move the other way. In the condition that hid AI answers, none of these measures deteriorated significantly. Removing the AI summary did not make the search experience worse.
The two conditions left differently shaped marks. Under AI Mode, clicks, sessions, trust, and satisfaction all moved down together. Under hidden AI answers, only external click-through moved up, and the perception measures were indistinguishable from the control. One condition pulled several metrics down at once. The other returned traffic without touching the user experience.
The table below lines up both conditions against the same set of measures.
| Measure (vs. control) | AI Mode enforced | AI answers removed |
|---|---|---|
| External click-through | −18.8pp | +8.8pp |
| Share of users clicking news sites | −12.5pp | Increase, not robust |
| Search sessions per day | −0.92 | No significant change |
| Minutes per session | +0.43 min | −0.59 min |
| Trust in information on Google (7 pts) | −0.34 pts | No significant change |
| Satisfaction | −0.73 sd | No significant change |
Figures for the AI answers removed condition are LATE estimates because of low compliance. Source: arXiv:2608.18352, Tables S1 to S5
The open-ended responses pointed the same way. Asked for their overall impression after a week of AI Mode, the 309 participants in that arm were more often negative than positive, 33.6% against 29.3%. The most common complaint was not accuracy but control. Loss of control and agency appeared in 17.6% of responses, difficulty reaching a specific website in 15.3%, and limited links or source diversity in 13.4%. On the positive side, efficiency and time saving led at 14.0%.
The remaining codes sharpen where the dissatisfaction sits. Responses calling AI outputs too verbose came to 6.2%, and those raising hallucination or accuracy concerns to 4.6%. Factual error, the first charge usually leveled at AI search, was mentioned far less often than the feeling of not choosing where to go next.
Conditions to read this with: the sample is US-based Chrome users, younger, more educated, and more left-leaning than the population. The observation window is seven days of treatment, so long-run adaptation is out of reach. Most of all, participants used AI Mode for 0.6% of their searches before the experiment began. That arm measured a forced, full-adoption scenario rather than voluntary use. It is closer to an upper bound on what happens if Google makes AI Mode the default.
In numbers, 75% of the sample was under 45 and 87% had at least some college education. One more condition applies to the timing. These values measured AI Mode before advertising arrived. AI Mode surfaced no ads during the study, and the authors expect user satisfaction to move away from this pre-monetization baseline once Google decides how to monetize it.
If answers keep scaling, who writes the source?
Filing this away as a search marketing result reads only half of it. AI answers are built from documents that someone wrote, used once as training corpus and again as grounding material at query time. If the revenue of the people writing those documents is tied to traffic, then every increment of the synthesis layer trims the incentive to keep producing the source layer. This experiment put a causal estimate on the first step of that chain.
Observations pointing the same way had already accumulated. Prior work cited in the paper found that AI Overviews reduced English Wikipedia traffic by 15%, and separate analyses found that visits and questions on Stack Overflow both declined after ChatGPT arrived, with no corresponding drop in post quality, which the authors read as genuine displacement rather than a culling of weak content. What this paper adds is a controlled replication of those observations, plus evidence that the decline is not confined to Wikipedia or news.
Regulators are pressing on the same point. In their discussion the authors cite the German case. According to Reuters, Germany's media regulator concluded that Google's AI Overviews amount to the company's own content rather than a mere display of third-party information, which brings them under German media law. The distinction between an intermediary and a publisher sits on exactly the axis this experiment measured: whether search sends people outward, or holds them with an answer.
For anyone working with data, the experiment leaves two practical questions. The first is whether we can keep assuming that the web corpus we ground on will be refreshed at its current rate. The second is what the fallback looks like if that assumption weakens, whether licensing deals, first-party collection, or nothing yet. How long it takes a traffic decline to turn into a content decline is something nobody has measured.
Editor's Note: this is why Pebblous looks at provenance and refresh cadence together when it talks about AI-Ready Data. Model performance is measured on a snapshot taken at some moment, but whether new snapshots can keep being taken depends on the economics of the people producing the data. The rate at which the upstream of a supply chain dries out does not show up in model metrics right away.
References
Primary Source
- 1.Wang, S. T., Gleason, J., Bart, Y., Wilson, C., & Metaxa, D. (2026). "AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence." arXiv preprint.
- 2.Agarwal, S., & Sen, A. (2026). "The Impact of Google AI Overviews on Publisher Traffic and User Experience: Evidence from a Field Experiment." SSRN Working Paper. A separate study from this article's primary source — see FAQ for the distinction.
Academic Papers
- 3.Khosravi, M., & Yoganarasimhan, H. (2026). "Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia." SSRN Working Paper 6164926.
- 4.Burtch, G., Lee, D., & Chen, Z. (2024). "The Consequences of Generative AI for Online Knowledge Communities." Scientific Reports, 14(1).
- 5.del Rio-Chanona, R. M., Laurentsyeva, N., & Wachs, J. (2024). "Large Language Models Reduce Public Knowledge Sharing on Online Q&A Platforms." PNAS Nexus, 3(9).
Industry & Regulatory Sources
- 6.Reuters. (2026-07-14). "German Media Regulator Says Google's AI Overviews Subject to German Media Law."