Executive Summary
The debate over AI training data has long been stuck between two extremes: free scraping that anyone can help themselves to, or strong copyright that protects creators tightly. A recent working paper from an MIT economics collaboration uses a game-theory model to argue that both extremes fail. The sharper twist lies on the strong-copyright side: how does a right strengthened to protect creators end up pushing them out first?
What the model produces is paradoxical. The stronger copyright becomes, the sooner the most original creators cut their effort. The authors call this the originality penalty. When an AI company stands as the single large buyer, it has the greatest incentive to bargain down the price of irreplaceable, original content. The very rule meant to protect creators pushes out the people it was supposed to protect first.
The authors' conclusion is that copyright should not be treated as a binary; the market itself needs to be redesigned. Their proposed alternative is a data intermediary that represents creators, bargains collectively with AI companies, and returns a larger share to original contributions. Because these are the results of a pure theoretical model, they are best read for their intuitions and implications rather than for the equations behind them.
The three figures below are the core results the paper draws from its game-theory model. Read them as equilibrium values produced by a model, not as measured empirical data.
about ½
Innovator's equilibrium effort
Falls to roughly half of the social optimum
about 0.79
Follower's equilibrium effort
A much smaller loss than the innovator's
finite ceiling
Future model precision
Capped once original creators drop out
The common sense that stronger copyright protects creators
Once it became clear that generative AI trains on text, images, and code made by humans, policy discussion split roughly into two camps. One side sees training as an extension of fair use and defends free access. The other holds that protecting creators means making copyright stronger to block unauthorized training. This binary is not confined to armchair debate — it is already being fought in court (What may an LLM learn from? Europe's top court asks the first question). Intuitively, the second side looks like the creators' side: the stronger the right, the better off the creator ought to be.
A working paper co-authored by researchers at the MIT Operations Research Center, MIT Sloan, and the University of Washington School of Law — "Market Design for AI: Beyond the Copyright Binary" — challenges that common sense head-on. The authors set up free scraping and strong copyright as separate game-theory models, and show that both extremes end in market failure. Free scraping gives creators no compensation at all, while strong copyright actually weakens the incentive to create.
Strong copyright fails not because creators go unpaid. The problem is on the other side of the market. On the buying side of training data, an AI company stands as effectively the only large buyer. In a structure with many sellers and something close to a single buyer, the buyer holds the power to push prices down. Even if copyright grows strong enough to attach a price to each individual work, the power sits with the side that bargains the price down, not the side that names it. That is where this model begins.
The most original creators are punished first
The static model the authors build runs in a set order: the AI company first offers a price to each creator, and the creators, seeing that price, each decide how much effort to put in. Here creators fall into two types — original contributors who do not follow the crowd, and followers who echo an already widespread trend. The original contributor's content carries information that cannot be found in similar form elsewhere; the follower's content carries information that overlaps with everyone else's.
The model's central result comes from here. In equilibrium, the original contributor puts in only about half the socially desirable level of effort, while the follower's effort falls far less. The authors name this asymmetry the originality penalty. The person producing the most valuable data receives the harshest punishment.
The intuition runs like this. As the single large buyer, the AI company wants to hold prices down to save cost. The follower's content already overlaps and has low marginal value, so there is little to bargain away. Original content, by contrast, is irreplaceable and therefore expensive — which is exactly why cutting purchases of it saves the most money. So as copyright grows stronger, it is the original creators' purchases that get cut most heavily, and they are the first to stop putting in effort or to be pushed out of the market.
This conclusion dovetails precisely with an observation Pebblous has covered before: in the AI content licensing market, only a handful of brands with bargaining power get a price, while most publishers remain free training data (The AI content licensing market where leverage sets the price). What that article saw in the field, this paper explains structurally, through game theory, as to why it happens.
The better AI gets, the worse the data gets
The authors extend this insight into a dynamic model that unfolds over time, and a second market failure emerges. As the model's precision rises — that is, as AI gets better — people lean more on the AI's drafts than on making something new by their own effort. As a result, newly produced content grows more alike, and this homogenized data flows back into training.
The vicious cycle completes here. A good model makes people depend on AI, that dependence homogenizes content, and the homogenized data is fed back into training, making bias harder to filter out. The authors call this the curse of precision. Once original creators drop out entirely, the model concludes, no amount of piling up merely correlated data can lift future model performance past a finite ceiling.
The paper points to signals already being observed as evidence. After conversational AI spread, traffic to developer Q&A communities fell sharply, and AI summaries at the top of search results are eating into the visits that once went to publishers. When the consumption of AI answers replaces the space where people used to create and reference original work, the very supply of new originals for the next generation of models to learn from shrinks. The real crisis for AI-ready data may lie not in refining quality but, earlier than that, in the supply drying up.
The data intermediary: reward, not penalty
If both extremes fail, the remaining path is to redesign the market. The mechanism the authors propose is a data intermediary — a body that represents scattered creators and bargains collectively with AI companies as a single counterparty. When creators haggle over price individually, they are unilaterally overpowered by the single large buyer; bundled together, their bargaining creates a countervailing force to offset that power.
The heart of the design is separating the incentive to produce from the distribution of rewards. Creators are paid at their marginal cost of effort so the incentive to keep creating stays alive, and on top of that a fixed subsidy is layered in to guarantee participation. When that subsidy is divided, it is allocated in proportion to how much each creator actually contributed to the model's precision. As a result, the more original the contributor, the larger the share they receive. It turns the penalty that strong copyright placed on originality into a reward proportional to contribution.
This intermediary differs from existing copyright collectives such as ASCAP or BMI. Where those bodies existed to cut the transaction costs of scattered licensing, the intermediary in this model has a new function, the authors stress: building a countervailing force against the AI company's buying power. This logic of bargaining — negotiating on creators' behalf and pricing innovation — is already appearing in the language of labor, in recent collective bargaining that required actors and voice performers to negotiate with their unions before handing their data over for AI training.
This paper is a pure theoretical model, not a prescription tested in the field. Yet the question it raises cuts straight to the heart of practice. Laying down a settlement ledger to compute who gets paid how much is a question of how to divide data that already exists. This paper asks the step before that: how to design the market so that people keep making good data in the first place. Before you finely compute how to split the pie, the question is whether you can keep the pie being baked at all.
References
Academic papers
- 1.Dai, Y., Farboodi, M., Golrezaei, N., & Shahshahani, S. (2026). "Market Design for AI: Beyond the Copyright Binary." arXiv:2606.12260.