Executive Summary

A paper posted to arXiv on August 7, 2026 argues for moving watermarks to a different job. Daniel Susser, John Thickstun, and Gili Vidan propose that instead of using the signal as a detector that rules on whether a given piece of content came from a machine, we use it as an instrument that measures how much synthetic content is circulating in the media ecosystem. The model they reach for is epidemiology, where testing wastewater reads the spread of infection in a community without interrogating anyone's medical history.

The argument does not turn on pushing accuracy higher. It turns on the requirements themselves changing. A verdict on an individual makes one false positive into one wronged person, so it demands something close to perfect accuracy, while an aggregate indicator can read a trend from a weak signal. In exchange, a different set of chores appears: a representative sample, a time series, and agreement on what is being counted.

Manufacturing has already made this move once, from inspecting every finished unit to reading the state of a process off samples and a time series. That analogy is not the paper's language, though. It is ours, added from the data quality side, and we have written down where it breaks.

Key Numbers

The first two numbers show the damage and the contamination synthetic content has already produced. The last one is what this paper does not yet have.

Sources: Nature, U.S. Attorney's Office, SDNY, arXiv:2608.07337

12x

Rise in fabricated citations

Nature reported an audit finding that they went from 1 in 2,828 papers in 2023 to 1 in 277 by early 2026

$10M

AI song streaming fraud

Royalties siphoned off by inflating play counts with bot accounts, charged in 2024 and pleaded guilty in 2026

0

Measurement pilots in the paper

What the authors offer is a conceptual redesign rather than a quantitative experiment

1

Watermarks Break Down on Individual Verdicts

The paper opens with a question from a 2018 New York Magazine piece. How much of the internet is fake? Max Read's answer was that a lot of it is, and the turning point he identified was not the volume of fakery but an inversion of the baseline. Content that no person made starts to serve as the reference point, and human behavior gets evaluated against it for authenticity. The authors begin there to show where the expectations loaded onto detection came from. The more the category of real wobbles, the harder people look for a tool that will hand it back, and watermarking was placed in that slot.

A watermark is not a visible stamp. It is a statistical signal. The basic form of language model watermarking tilts the sampling probabilities during generation so that certain groups of tokens come up slightly more often, then tests later for whether that bias survives. What the detector returns is a probability rather than a verdict, and wherever the threshold is drawn, false positives and false negatives move together. No threshold drives both to zero.

Move the threshold and false positives / negatives trade off together Threshold Human-written text AI-generated text Move left: FN↓ FP↑ Move right: FP↓ FN↑ Schematic of the argument in arXiv:2608.07337 (conceptual, not a measured distribution) | Original Pebblous diagram
▲ No single threshold drives both false positives and false negatives to zero | Original Pebblous diagram

Text has it worse. Unlike images or audio, sentences carry little spare capacity. Short pieces, formulaic answers, and paragraphs dense with numbers and citations leave less room to carve a signal into, and what remains fades under light paraphrasing. Theoretical results point the same way. An impossibility argument published in 2024 showed that a generic attack can strip a watermark without much loss of quality, and a separate line of work sets out why a watermark that is both publicly verifiable and robust is so hard to build.

Then there is the question of publishing the detector. The detection dilemma, framed in 2021, states the tension plainly. The wider a detection tool is opened, the more useful it becomes to journalists and citizens, and the more easily anyone trying to evade it learns to route around it. As long as the goal is catching individual pieces of content, releasing the tool weakens the goal.

Last comes the person reading the signal. Verdicts land in courtrooms, school disciplinary processes, and platform suspension notices, and those are not places where a probability stays a probability. Ninety percent reads as guilty, and the cost of the remaining ten percent is paid alone by whoever was misclassified. That burden is not evenly spread either. As a 2020 study of speech recognition accuracy across speaker groups showed, the errors of statistical classifiers tend to pile up on minority groups. A judgment that something looks machine-written can work against non-native speakers in much the same way.

The paper's point is not that the technology is still immature. Asking whether one piece of content was made by AI demands a level of certainty that a statistical signal cannot carry. The demand is excessive, which is not the same as the signal being useless, and that distinction is where the next section begins.

2

Change the Question, Change the Requirements

The paper splits watermark use into two branches. One is forensic, establishing the provenance of an individual item. The other is measurement, gauging what share of a market or a platform is synthetic. Both use the same signal, and the conditions they demand barely overlap.

Requirement As individual detection As ecosystem measurement
Accuracy One false positive wrongs a person, so it has to be near perfect Weak signals pooled together are enough to read a trend
Robustness Has to survive every attempt to strip it Even partial resistance keeps a sample, as long as stripping takes effort
Unforgeability Evidence in a proceeding needs a cryptographic guarantee Some forgery shifts the distribution without collapsing the trend
Who interprets it Courts and general readers have to read probabilities correctly Analysts at platforms and institutions handle it as statistics
Point of intervention Blocking and takedown before publication Adjusting recommendation weights and payout rules after the fact
Coverage gaps One miss is a failure Survivable as long as the sample stays representative

Compiled by Pebblous from the argument in arXiv:2608.07337.

Nothing in the right-hand column says accuracy stops mattering. What changes is the size of the error that can be absorbed. In individual adjudication, a 5% false positive rate means one person in twenty is wronged. In a measurement of prevalence, the same 5% is a bias that can be corrected or noted. So instead of lengthening the list of things watermarks cannot do, the paper asks again what this signal already does well.

The detection dilemma from the previous section also cools down here. The tension between opening a tool up and making evasion easier is sharpest when the purpose is catching individual pieces. When the purpose is measuring a share, some traffic slipping out of the sample still leaves enough behind to read a trend, so publishing the tool costs that much less. This is the basis for the authors' claim that moving to measurement does not dissolve the governance homework, but does leave a more tractable version of it.

3

Factories Already Made This Move in the 20th Century

This section is our reading. The paper does not reach for factory quality control. Its analogy is epidemiology and wastewater surveillance. Seen from the data quality side, though, the proposal has the same shape as a road manufacturing already traveled once in the twentieth century.

Early quality control defaulted to inspecting every unit. Finished products were examined one by one and defects pulled out. The method holds only on the premise that the inspector's eye is perfect, and it collapses as volume grows. Statistical process control took its place. Instead of looking at every unit, you sample and measure, plot the variation as a time series, and adjust the process when the flow crosses a control limit. Accuracy on any single unit went down, while visibility into the state of the whole process went up.

From Inspecting Every Unit to Reading the Process 100% Inspection Judge each unit alone Only works if the inspector is perfect Collapses as volume grows Shift Statistical Process Control Sample and track over time Adjust the process when limits are crossed Less accurate per unit, clearer view of the whole Read against manufacturing quality-control history — a Pebblous reading, not the paper's own analogy | Original Pebblous diagram
▲ 100% inspection depends on a perfect inspector; statistical process control reads the trend from a sample and a time series | Original Pebblous diagram

The move the paper proposes travels the same distance. Used as a defect inspector, a weak watermark signal reads as failure. Used as a process indicator, the same signal becomes a gauge that reports a trend. A gauge is allowed to be off by a tick. What it needs is for the ticks to be marked the same way every time.

Where the analogy breaks is just as clear. A factory controls its own process and knows exactly where to draw its samples. The information ecosystem has no controller. Models that embed nothing, open-weight models people run themselves, and participants who strip the signal on purpose all sit outside the sample, and nobody knows how large that outside is. Above all, a factory has a normal range, while the share of synthetic content has no baseline at all. Nothing yet tells us whether 30% is an alarming number or an ordinary one. So the first task in moving to measurement is not a measurement technology. It is agreeing on what to count.

4

What Happened on Music Charts and in Paper Review

The first setting the paper works through is music streaming. In September 2024 the U.S. Attorney's Office for the Southern District of New York charged a music producer with siphoning off more than $10 million in royalties by streaming hundreds of thousands of AI-generated tracks through bot accounts, and the case closed with a guilty plea in 2026. Those were fully synthetic tracks pushed in bulk, so the signal was clear. What is far more common in the market is a human-made track with generative tools layered onto parts of it. Those leave only a weak signal, and a weak signal is thin ground for pulling a specific track or withholding a payout.

In measurement mode the same signal does different work. If a platform published the synthetic share by genre and chart band on a regular schedule, it would have grounds to adjust recommendation weights and payout rules without passing judgment on any single track. How that share moves quarter to quarter inside the top 100 is more useful for running a market than any one verdict.

The second setting is academic publishing, and as it happens, arXiv itself, where this paper was posted. On October 31, 2025 arXiv announced that in computer science it would accept survey articles and position papers only after they had cleared peer review. The reason was a flood of manuscripts turned out quickly with generative tools, more than moderators could handle. In 2026 word followed of a policy barring authors for a year when unchecked AI use is confirmed, with fabricated references as the clearest evidence.

Fabricated references function here as a naturally occurring watermark. A citation to something that does not exist is a trace of using a tool without checking it. The question is what to do with the trace. A submission ban is a binary judgment on a person, and the cost of misclassification falls entirely on the author. Stack the same signal into a time series by field and a different option opens up. Publish how fast it is rising and where, then tune review procedures, submission guidelines, and author training to that trend. The audit figures Nature reported, with fabricated citations going from 1 in 2,828 papers in 2023 to 1 in 277 by early 2026, are the product of exactly that kind of measurement rather than of catching papers one at a time.

Both cases have the same structure. Detection mode reaches a conclusion about an individual from an imperfect signal and hands that individual the cost of the error. Measurement mode pools the same signal into grounds for adjusting policy. The quality of the signal does not change. What changes is who pays for its limits.

5

What Measurement Would Require

Move the purpose and the homework moves with it. In place of a perfect detector, what the job needs is the basic craft of statistical surveying. The paper's conditions come to four.

  • A representative sample. Which platforms, languages, and formats go into the sample decides the result. Count only English text and the number describes English text, not an information ecosystem.
  • Time-series infrastructure. A single survey gives a number but not a trend. Measurement has to repeat under the same definition, and the moment the definition changes, comparison with earlier periods breaks.
  • Agreement on what to count. Is a wholly generated piece the same item as one whose draft was generated, or one where a few sentences were polished? Without a line here, two institutions publishing figures that contradict each other becomes routine.
  • A form of publication. The place the paper points to is the artifact a platform's internal audit produces, meaning a transparency report in a form that can leave the building. Publishing enforcement statistics on a schedule is already common practice, so the synthetic share can ride in the same document.
Four Conditions for Building a Measurement Regime 1 Representative sample Platform, language, format shape the result 2 Time-series infra Repeat measurement under one definition 3 Counting agreement What counts as one item Heaviest 4 Publication format As a transparency report, on schedule The four conditions arXiv:2608.07337 lays out | Original Pebblous diagram
▲ In place of a perfect detector, the basic craft of statistical surveying — the heaviest of the four is agreeing on what to count | Original Pebblous diagram

The quietly heaviest of the four is the third. The U.S. National Institute of Standards and Technology defines synthetic content broadly, as information significantly altered or generated by algorithms. The definition deliberately draws no line between wholly generated and partly edited, which lowers the risk of leaving regulated material out, but it leaves the unit of counting empty. Anyone building a measurement regime has to draw that line themselves, and the line is what fixes the meaning of the numbers they publish.

Regulation is moving on a parallel track. Article 50 of the EU AI Act requires machine-readable marking from August 2, 2026, and California's SB 942 took effect on January 1, 2026. Both tell providers to attach a mark, and both leave open what the mark should be used to count. If our earlier piece On August 2, Machine-Made Text Gets a Tag covered the duty to attach the tag, what this paper aims at is the next column over, which is what the tag is for. Set it beside The Economics of Synthetic Data Contamination, which read watermarks as a price signal, and the same mark turns out to have three distinct institutional uses.

The authors admit to weaknesses of their own. The largest is complacency. Watermarks are cheap and easy to attach, so policymakers may mistake them for the solution and defer harder questions about copyright, labor, and platform responsibility. Coverage gaps outside the sample and the possibility of forgery remain. Watermarks that identify individual accounts carry a real risk of becoming surveillance tools, which the authors themselves do not recommend. And the paper has no pilot measurement data to test the proposal against. It reframes a concept rather than reporting results.

6

The Same Question Waits in Our Own Data

Data quality work is full of signals discarded for insufficient accuracy. A dedup classifier that misses twice in ten tries gets pulled out of automated cleanup, a label review model with low confidence has its output ignored, and an outlier detector that alerts too often gets its notifications switched off. Judged as tools for reaching conclusions about individual records, all of those calls are reasonable. Put the same signals on a chart of ratios over time and the story changes. The fact that the share of suspected duplicates doubled since last month tells you something upstream in the pipeline changed, even if you cannot say which records are duplicates.

Three questions come before deciding whether to discard a signal or accumulate it.

  • Is this signal being used for individual adjudication, or read as a ratio? Excessive accuracy demands usually live on the first side.
  • Is it measured repeatedly under the same definition? Change a threshold quietly and the trend breaks that day.
  • Does the team give the same answer for what counts as one item? Split definitions put different numbers on two dashboards.

Editor's Note: What Pebblous runs into most often in data quality work is not an absence of usable signals. It is a stock of signals already thrown away for falling short of a verdict. This paper says the same thing about the information ecosystem. Making a signal more accurate and deciding what to count with the signal we already have are two different jobs, and the second one has mostly not started.

R

References

Academic Papers

Industry & Press

Official Documents