Executive Summary

Researchers at the Centre for the Governance of AI (GovAI) and the Hertie School swept government services in twelve countries and pulled out 84 cases, spread across eleven jurisdictions, where requests surged because of AI agents. They call the phenomenon agentic flooding. It is not a warning about what might happen; it is a count of paperwork that has already been filed.

The part that stays with you is what governments did next. In 14 cases they pushed back by adding friction, and the measures they reached for were the ones a government can deploy today, such as reinstating a fee or blocking an IP range. The paper notes that these measures screen out poorer and less digitally literate users first. The technology meant to widen the door is now narrowing it.

South Korea is in the sample. Applications for payment orders in the civil electronic litigation system appear as one case, cited in a passage about people with low legal literacy now being able to file suit.

Key Numbers

Source: Schmitz et al., arXiv:2608.16603 (2026-08-17)

2,288 → 84

candidate services that cleared the bar

A case had to show a cost-reduction pathway, evidence of changed demand, and an explicit attribution to AI

87%

share held by a single mechanism

Language models generating text cheaply and in bulk; sophisticated behavior such as autonomous browsing has not been observed yet

4,000 pages

a single letter filed with a German social court

The volume grew inside one filing rather than across many, which is exactly what submission caps fail to catch

14 cases

responses that added friction

17% of the 84, covering reinstated fees, IP blocks, and limits on filing through intermediaries

1

2,288 Services Screened, 84 Left

The researchers started with 2,288 candidate government services across twelve countries, roughly 190 per country, and kept only the ones that cleared three conditions at once. There had to be a plausible pathway by which AI lowered the cost of using the service, evidence that demand patterns had actually shifted, and an explicit statement from the government or a credible third party attributing that shift to AI. Fewer than one in twenty survived, leaving 84.

Twelve countries were swept: Australia, Brazil, Denmark, Estonia, France, Germany, Japan, the Netherlands, Singapore, South Korea, the United Kingdom, and the United States. Only eleven of them produced a case that met the bar, which means one country came back empty. Cases were collected through LLM-assisted qualitative coding, split into a broad discovery pass, an iteration pass that tightened the criteria, and a finalization pass, with a human checking the work at every stage.

Fewer than one in twenty survived 2,288 candidate services (12 countries) ① Is there a plausible cost-reduction pathway? Services with a confirmed pathway ② Did demand patterns actually shift? Services with confirmed demand shift ③ Did government or a third party name AI explicitly? 84 cases Only 11 of the 12 countries produced a case One country cleared none of the three criteria
▲ The three-stage screening from arXiv:2608.16603, redrawn as a funnel | Original Pebblous diagram

How strong the evidence is varies, and that is worth weighing too. In 58 of the 84 cases a government official pointed at AI directly; the remaining 26 rest on third parties such as news outlets or legal trade press. By domain, judicial and legal services led with 19 cases, followed by regulatory complaints at 10 and welfare and social security at 9. Public services and judicial services together accounted for 73% of the set, and the rest were participation channels more than services, things like consultation submissions and freedom of information requests.

The shape of this is not new. Public administration scholarship already had a term for the state in which demand outruns processing capacity and the counter jams up, the congested bureaucracy, and researchers had already observed a marked rise in self-representation in United States federal courts. What the 84 cases do is fill that theory in with numbers and the names of jurisdictions.

The paper then takes a step back. The authors state plainly that 84 cases neither prove causation nor represent government services as a whole. There is still no data that can say what share of a rise in filings AI produced. What this sample offers is an existence proof: this is happening, and at this scale.

2

It Is Not Just the Count

Agentic flooding, as the researchers define it, is a state in which the volume or complexity of requests reaching a government service surges because of interactions with AI agents, to the point of putting real strain on processing capacity. One more condition is attached: the surge has to trace back to agents lowering the cost of interacting. It is a device for separating a rise in filings from a rise in filings caused by AI.

The surge shows up along two dimensions. Quantitative flooding, where the number of requests grows, appears in 50 cases; qualitative flooding, where a single request grows in length and complexity, appears in 76; and 42 cases show both. A letter running past 4,000 pages arrived at a German social court. The distinction matters operationally because it splits the available responses. Capping daily submissions, a common move, does nothing about qualitative flooding. One person may file one document, and if that document is 4,000 pages, the burden is unchanged.

Growth along two dimensions Quantitative More requests filed 50 cases Qualitative Longer, more complex filings 76 cases 42 cases both at once German social court: a 4,000-page letter Capping submission counts does nothing for qualitative flooding
▲ The split between quantitative and qualitative flooding across the 84 cases in arXiv:2608.16603 | Original Pebblous diagram

Eighty-seven percent of the observed mechanisms converge on a single one: language models generating text cheaply and in bulk. Agents navigating agency websites on their own and working through procedures do not appear in the sample yet. The authors offer two readings. Text generation is the most mature and most accessible capability, and text submissions are also the ones where anomalies stand out. A filing that arrives through a structured form is hard to attribute to a human or an agent, so it can be happening already without ever reaching the sample.

3

Lucrative and Complicated Desks Break First

The paper offers a matrix that measures the risk to each service across 13 factors. What the table shows most clearly is the spot where two conditions overlap. Services that pay well when an application succeeds, and where procedural complexity or a demand for expertise has been holding demand down, are the first to give way. Tax filing, social court litigation, property value appeals, and small claims all sit in that overlap.

The matrix measures along two axes. One is how likely flooding is: how mature agent capabilities are and how cheap they are to use, how exposed the intake channel is and how much effort a submission takes, and how much a successful application is worth. The other is how severe it would be if it happened: the effort of handling one case, how close to capacity the service already runs, whether budget and staffing can be expanded at all, and whether statutory processing deadlines apply. Most of those 13 factors are values each agency already knows, independent of AI. That is why the table can be filled in today, without anyone having to predict how quickly AI advances.

Put the other way around, complexity has been serving as a defensive wall. Nobody designed it that way, but that was the effect. The paper splits administrative burden into three parts: learning costs, spent understanding the rules; compliance costs, spent filling in forms and assembling documents; and psychological costs, spent dealing with a bureaucracy over and over. Agents that summarize, stitch context together, and carry memory shave all three at once. Commercial incentives ride on top of that. Claims management firms, which file standardized volume claims such as flight delay compensation on behalf of clients and take a cut, are already putting AI in their advertising.

Desks where complexity was the wall Procedural complexity (low → high) Payoff if successful Highest-risk zone Tax filing Social court claims Property value appeals Small claims Certificate issuance Change of address Procedures that paid well and were hard to navigate give way first Complexity was a wall nobody designed, and agents clear it at no cost
▲ The central combination from the risk matrix in arXiv:2608.16603, redrawn as a concept diagram | Original Pebblous diagram

Here is part of the case table the paper published. South Korea appears through applications for payment orders in the civil electronic litigation system, classified as qualitative flooding. The count of filings is not what grew; the review burden inside a single filing is.

Jurisdiction Service Type
Australia Federal freedom of information requests Quantitative
Germany Social court litigation Quantitative + qualitative
United Kingdom Money claims online Quantitative + qualitative
Japan Strategic Energy Plan public comments Quantitative + qualitative
Netherlands Municipal property value appeals Quantitative + qualitative
South Korea Civil electronic payment orders Qualitative
Brazil Temporary disability benefit claims Qualitative

The Brazilian case has a different texture. Attempts to claim temporary disability benefits with AI-generated fake medical certificates were mixed in. Most of the rest are people with a valid entitlement exercising it. The reason the paper bothers to keep that distinction is that tools built to catch fraudulent claims cannot address a surge made of legitimate ones.

4

The Fastest Response Is the Riskiest One

In 56% of the 84 cases the government moved in some fashion, though a good share of that amounted to issuing non-binding guidance. Measures that actually reduced incoming requests appear in 14 cases, and the method was to add friction. Australia brought back a fee for freedom of information requests, Japan blocked specific IP addresses from a public comment channel, and the Netherlands stopped firms from taking a cut for filing property value appeals on someone's behalf.

The reason friction comes first is simple. If a payment system is already in place, a fee can go up today, and even where a legal change is needed the surface being touched is narrow. Expanding processing capacity, by contrast, runs through budget cycles, organizational restructuring, and sometimes legislation.

There is one more way to reduce incoming requests: reducing what a successful application delivers. Cut the benefit or narrow eligibility and the reason to apply shrinks with it. That lever, though, is not something an executive agency decides alone, and because it runs through legislation it has not become the card anyone reaches for in a hurry.

The fast response and the slow one Adding friction Fees, IP blocks 14 of 84 cases Deployable today Screens out the users who needed help most Expanding capacity Redesign, agent deployment 13 redesigns · 21 deployments Needs budget, legislation, restructuring Preserves access, but arrives after the surge Now Months Years Decades Time to deploy The card you can play first is not the card you should play The slow option only exists for those who started before the crisis
▲ Deployment speed and access impact of adding friction versus expanding capacity | Original Pebblous diagram

The problem is who friction stops first. Fees and identity checks look as though they apply equally to everyone filing a request, but the people actually screened out are the ones without money, without digital literacy, and with the most urgent need in the first place. The paper reads this as a path toward procedural inequality, eroded access to justice, and lower trust in government. Tools that fill in complicated forms on someone's behalf were supposed to lower the wall for exactly these people.

Friction may not hold for long either. As CAPTCHA showed, rising agent capability can render the friction itself useless, leaving a threshold that only trips humans. There are also plenty of areas, welfare services among them, where charging a fee is prohibited by law. The authors connect these stopgaps to what public administration scholarship calls muddling through.

5

Why the Frictionless Road Is Still Unbuilt

The route that does not harm access has begun as well. Services were redesigned in 13 cases, and agents were put into the processing pipeline in 21. The United Kingdom grouped comments submitted to local planning consultations using sentiment analysis, Brazil strengthened detection of AI-generated fake medical certificates, and courts in South Korea introduced a service that verifies whether a case number is genuine.

Most of these, though, are limited tools that handle a single step. Structural overhaul takes years to decades, and in the meantime budgets and staffing are built around an assumption of constrained demand. That is where the ordering of time becomes the problem. Expanding capacity is not a choice that remains within reach once a surge has started. It is a card only the organizations that began before the crisis get to play.

The three things the paper names as doable right now are these.

  • Audit exposure. The risk matrix lets an agency rank its own service portfolio today. It is the one measure that requires no forecast of how quickly AI advances.
  • Build a digital identity strategy. Electronic authentication makes per-claimant request limits automatically enforceable, and prefilling existing data lowers the burden on ordinary users at the same time. More than 100 jurisdictions already have the infrastructure, though maturity varies widely.
  • Resolve legal uncertainty. Whether restricting submission channels or raising fees is lawful remains unclear in many jurisdictions. Finishing that legal review in advance is what sets response speed during a crisis.

The tone the authors take in their conclusion is not catastrophic. They put the odds of an acute administrative collapse low, while writing that budget adjustments and degraded user experience look hard to avoid in at least some cases. What they do warn about is friction hardening into the default response, which would seal off the very gains in access that AI promised. The preparation needed now, on the other hand, points in the same direction as reforms digital government researchers have been asking for all along.

For anyone working with data, the question this paper leaves behind applies well outside public administration. When the cost of arriving at any intake point suddenly drops to near zero, you find out whether what was protecting that intake point was policy or merely inconvenience. If it was inconvenience, something has to be built in its place, and keeping that something from becoming a threshold is the task right after.

Editor's Note: We see a similar scene in data pipelines at Pebblous. Put the filter in after the volume arrives, and what gets caught is usually not bad data but unfamiliar formats. Validating upstream has always been cheaper.

The original paper is available at arXiv:2608.16603.