Executive Summary

When technology shakes up work, politics turns to compensation. That is the sequence political economy has assumed for decades. Two researchers based at Amsterdam's two universities went through plenary speech in 33 national parliaments, pulled out the 5,317 speeches that connected AI to work, and counted what those speeches proposed doing about it. The sequence did not show up. Compensation accounted for 2.3 percent of response-frame mentions. This report looks at how that empty box connects to the question of what the people who make training data are owed.

The axis of conflict in parliament was not how to repair the damage but whether to push this technology forward or hold it back. The authors call it the enablement–regulation axis. Open the regulation side, though, and the largest thing inside it is not algorithmic management at work but copyright and protection of the creative sector. When politics talks about the people who made the data, the vocabulary it holds is ownership, not payment. That internal breakdown comes from an exploratory analysis the authors added as an interpretive aid, so it is worth reading for direction only.

For anyone designing a data supply chain, the useful reading here is not a regulatory forecast but a schedule. Nothing in this evidence supports a plan that waits for legislatures to write the rules on payment. The authors themselves state that their evidence cannot establish a representation gap. What is clear is what is currently not being said on the political stage.

2.3%

Compensation's share

The denominator is all response-frame mentions drawn from 5,317 AI-work speeches

55.2%

Enablement and investment's share

First on the same denominator. Regulation follows at 21.8%, training at 20.6%

36.1%

Copyright's share inside the regulation frame

Denominator is regulation-frame mentions. Exploratory, not validated coding

125

Compensation mentions in absolute terms

Both compensation topics combined, across 33 parliaments from January 2023 to April 2026

1

The politics the literature predicted did not show up

There is a long-standing answer to what politics does when technology unsettles work. Automation cuts jobs, presses wages, and hollows out occupations, and politics begins over how to make good on the damage. Compensation and social insurance, redistribution and retraining were the items on that agenda. Section 2 of this paper condenses the tradition into one line: politics enters the story downstream of technology, only once exposure has occurred and its consequences for particular workers are becoming visible.

The same section also records what that tradition misses. Technologies do not enter labor markets as autonomous shocks with fixed consequences. The same technology can displace workers or augment them, intensify surveillance or support collective capacity, concentrate productivity gains or spread them. Which of those it becomes depends on how it is adopted and governed. If that is true, the question facing public authority is not only how to repair disruption. Whether to accelerate the change or constrain it sits alongside it.

The paper's second point runs straight into this report's own concern. What AI does to work does not stop at displacement; it rearranges supervision, monitoring, evaluation, expertise, and authority. In the paper's own words, a worker whose tasks are monitored by an algorithm, whose professional judgment is displaced by a model, or whose output is priced against automated alternatives experiences a shift in economic power well before any unemployment spell or benefit claim would register in a compensation-focused account. A framework centered on compensation cannot see that shift. The question of what data work is owed, which this report reaches later, sits in exactly that position.

The paper that went looking for the predicted politics in actual legislatures was posted to arXiv on 2 September 2026. It is titled "Meeting the Coming Wave: The Emerging Politics of AI and Work across 33 Parliaments," by Juliana Chueri of the Department of Political Science and Public Administration at Vrije Universiteit Amsterdam and Petter Törnberg of the Institute for Logic, Language and Computation at the University of Amsterdam. It is a version 1 preprint, and no peer-review status is stated.

What the two assembled is 1,514,950 substantive speeches delivered in 33 national parliaments between January 2023 and April 2026. Multilingual dictionaries retrieved a broad candidate set of AI speeches, and a second classification stage narrowed it to the 5,317 speeches that connected AI to work. Each of those 5,317 carries two codings: what the speech diagnoses AI as doing to work, and what it proposes doing about it. The second is the response frame, and it has four boxes. Regulation and restriction, enablement and investment, training, and compensation.

The four boxes were derived from theory rather than discovered in the data, and each answers a different political question. Table 1 of the paper sets out the mapping.

Response frame Political question it answers Examples the paper gives
Compensation How should losses or reduced labor demand be absorbed? Unemployment insurance, social protection, reduced working time, basic income
Training How should workers and organizations adapt? Reskilling, lifelong learning, AI literacy
Enablement and investment How should adoption and capacity be expanded? Firm support, infrastructure, public AI strategy
Regulation and restriction How should the pace and terms of deployment be governed? Worker voice, safeguards, limits on surveillance, moratoria, robot taxes

Right where this table sits, the authors note something in advance. Some boxes may fill up while others stay largely vacant, and that shape is itself evidence about how a new technological conflict is being constructed. That is why this report spends so long looking at an empty box. The list of examples on the bottom row, the regulation box, will be needed again in section 4.

1.1Measuring the four boxes

A single speech can receive one primary and one secondary response frame. The unit of this distribution is therefore not the speech but the mention. The four values below take all response-frame mentions drawn from the 5,317 AI-work speeches as their denominator, count primary and secondary tags together, and sum to 100 percent. This is where the paper's central result sits.

Response-frame mentions (base: 5,317 AI-work speeches, sums to 100%) Enablement and investment 55.2% Regulation and restriction 21.8% Training 20.6% Compensation 2.3% 0% 60%

Redrawn from the values in section 4.3 and Figure 7 of the paper. This is a distribution of mentions counting primary and secondary tags, with all response-frame mentions from the 5,317 AI-work speeches as the denominator. These four values come from the human-validated coding.

Parliaments barely talked about making good on the damage. They talked instead about what to do while the technology is still unfolding: accelerate adoption, prepare people and institutions, or govern deployment. In the phrase the authors use in their conclusion, what has emerged so far is not a politics of compensation.

1.2The surprise belongs to the literature, not the authors

This result is not a number that jumped out of the data at authors who were not expecting it. Section 2 of the paper sets out its expectations in advance, and the second of them points in exactly this direction.

"Because AI's labor-market consequences remain unsettled, AI-work politics should be organized primarily around whether adoption should be accelerated and enabled or conditioned and restricted, rather than around post-disruption welfare responses. Regulation and restriction, and enablement and investment, should therefore be more prominent than compensation."

So writing that the scholars were surprised by their own findings would be false. The distance opened up here is not between the authors' expectations and the data. It is between the data and the founding assumption of a literature that has explained automation politics for decades. That literature took politics to begin by cleaning up after technology, and parliaments were arguing about direction before the consequences had settled. The authors name the first a politics of compensation and the second a politics of adoption, and hold them apart.

1.3This debate is still small

Whatever shape the distribution has, the debate itself is still small. AI candidate speeches rose from 1.06 percent of all substantive speech to 1.70 percent, and the share of AI debate connected to work rose from 18.0 to 33.8 percent. Both figures contrast 2023 with 2026, and the 2026 values cover only the four months from January to April. Even after that growth, AI-work speech stays well below one percent of substantive parliamentary debate in 2026.

The growth is uneven and so is the geography. Speech has a different color from country to country. Spain, Chile, Belgium, and Italy are among the more threat-oriented cases, while the United States, Canada, Taiwan, and Germany lean strongly toward opportunity. One thing is nevertheless common. In most countries the largest primary response is enablement and investment, and compensation is marginal almost everywhere. The paper does warn against reading the national rankings literally. The sample was set by which countries have official records recoverable at speech level, and parliamentary institutions and transcript practices differ, so those differences can bleed into the ranking.

Taiwan's Legislative Yuan building — one of the parliaments where AI-work speech leans most strongly toward opportunity
▲ Taiwan's Legislative Yuan. The paper names this among the two parliaments with the highest density of AI-work speech, and its tone leans strongly toward opportunity. | Source: Wikimedia Commons

Growing fast while remaining small means, on the paper's reading, that parties are still finding their positions. The empty boxes this report examines are therefore a photograph of a moment rather than a permanent conclusion. The authors set out three ways the photograph could change: visible displacement could strengthen compensation politics, deepen contestation on the left, or create new opportunities for radical-right mobilization. Contracts that buy and sell data, though, are not waiting for the photograph to change. Clauses are being written somewhere right now.

2

How 1.51 million speeches became 5,317

Every ratio that follows in this report has passed through the process that narrowed 1.51 million speeches down to 5,317. What that process kept and what it threw away is what defines the range these numbers can be used over. Most of what a reader with the narrow interest of data labor will snag on is also settled inside that process, and section 5 rests on how it was settled.

2.1The sample was settled in two stages

Collection ran through country-specific parsers. Official APIs, XML feeds, HTML reports, Hansard-style transcripts, and PDF publications were each scraped and normalized into a common schema, with one row representing one substantive intervention. Procedural interventions and chair turns are already excluded at that point. The 1,514,950 speeches assembled that way are this study's full corpus. Two stages then produce the analytic sample.

Substantive speeches 1,514,950 · 33 lower chambers, Jan 2023 to Apr 2026 AI candidate speeches 19,411 · multilingual dictionary retrieval AI-work speeches 5,317 · the base for every share below Directly relevant 3,076 · robustness sample

Redrawn from the values in section 3 and Table 2 of the paper. The narrower 3,076-speech sample was used to check whether the results hold under a stricter definition; the response-frame distribution reported in this article is on the broad 5,317-speech sample.

The paper states the test that separates the two samples. A speech is directly relevant when AI, automation, or algorithmic systems are explicitly connected to employment, tasks, wages, skills, job quality, workplace control, labor-market institutions, or worker protection. It is indirectly relevant when the connection runs through something broader such as occupational transition, productivity, or public-sector capacity, and adjacent or irrelevant speeches were dropped entirely. Wages and worker protection are inside the conditions for direct relevance, so if a legislator had spoken on the floor about what labeling workers are paid, the speech would have been caught by this net. What box it would then have landed in is a separate matter, taken up in section 5.

The list of 33 countries runs from Argentina to the United States and mixes wealthy and middle-income economies, parliamentary and presidential systems. Kenya, the epicenter of the dispute over data labeling work, is on the list. So are Taiwan and Singapore, named as the two highest-salience cases. South Korea is not. The paper is explicit that this is not a probability sample but a set of countries whose records could be recovered at speech level.

2.2A language model did the coding, and humans contested it twice

Frame coding was handed to a language model in zero-shot mode. The runs used gemini-3.1-pro-preview on 6 and 7 July 2026 at temperature 0, and the annotator received the original-language speech text and the coding instructions only. It was not told the speaker's party family, government status, or the hypotheses the paper tests. Every classification had to be grounded in an explicit claim inside the speech.

Validation used 728 human-coded observations in two designs. Misreadings are common at this point. Two kappa values appear, and quoting only the better one makes the reliability look higher than it is.

Validation sample How it was drawn Relevance accuracy Kappa
Primary benchmark, 538 speeches General sample coded by the lead human coder 97.8% 0.947
Stress test, 190 speeches Deliberately enriched for boundary cases, earlier disagreements, and rare response categories 85.8% 0.679

The two values measure samples of different difficulty. The lower one was built by deliberately picking the places where the codebook was most likely to fail, and the authors state that it is not intended to estimate corpus-wide accuracy. Performance on the frame coding itself was measured among speeches that both human and model judged relevant: micro-F1 of 0.893 on diagnosis and 0.878 on response. That is a demanding bar, since in a multi-label task the complete set of assigned frames has to match.

Among the 114 speeches both coder and model classified as AI-work relevant, the direct-versus-indirect split agreed 96.5 percent of the time at a kappa of 0.916, with micro-F1 of 0.891 on diagnosis and 0.828 on response. Dividing speech into frames held up reasonably well even in the hard sample, in other words, and what dragged kappa down to 0.679 was the upstream judgment of whether a speech is about AI and work at all. In a sample assembled from boundary cases on purpose, that is the part designed to wobble.

2.3Two qualifications travel with the 2.3 percent

The headline figure of 2.3 percent carries two qualifications that belong in the body text with it. First, the value is on the broad 5,317-speech sample. Section 4.3 reports that the ordering holds when AI-work relevance is defined narrowly and only directly relevant speeches remain, and adds that in that sample compensation rises somewhat. How much is not stated in the body. A second robustness check sits in the same place: excluding Singapore and Taiwan, the two highest-salience cases, leaves the pattern unchanged.

Second, the place where that value should be is empty. The paper cites its appendix five times. Corpus construction details, the language-model prompt implementation, human validation details, the full clustering procedure, and the robustness checks just mentioned all belong to the appendix. No appendix is attached to the version 1 preprint on arXiv. That holds for both the HTML rendering and the 32-page PDF. This is not a matter of concealment but of a first preprint that has not yet acquired its appendix, and in the meantime the strongest sentence available is that compensation rises somewhat.

3

Parties split on diagnosis, not prescription

To find out why compensation is empty, look at where the parties actually divide. This paper's answer is that they divide on diagnosis rather than prescription. Parties say different things about what AI does to work, and that difference all but determines the prescription that follows. The values below come from a different table than the mention distribution in section 4.3. They are predicted shares from multilevel models with party-family and year fixed effects and random intercepts for country and country-party, and because diagnosis is also multi-label, threat and opportunity need not sum to 100 percent.

Party family Diagnoses AI as a threat Diagnoses AI as an opportunity
Radical left Roughly nine in ten Fewer than one in three
Greens Lean toward threat, though less decisively The paper gives no separate figure
Social democrats At a similar rate to opportunity At a similar rate to threat
Liberal, conservative, Christian democratic, radical right Around one-third 80 to 86%

The denominator is speeches with a clear diagnosis. What pairs with this table is that party families differ little in how often they talk about AI at all. Liberal parties have the highest predicted salience and greens and the radical left the lowest, but the uncertainty intervals overlap substantially. The contest is not over who raises the subject often. It is over what they call it when they do.

The coverage of that comparison is not the whole of it. Party affiliation was assigned at the date of each speech, producing 936 country-party units, and those were harmonized into seven families: radical right, conservative, liberal, Christian democratic, social democratic, green, and radical left. The paper states that parties which could not be mapped with sufficient confidence were excluded rather than forced into the scheme. As a result, party families are identified for 94.7 percent of AI-work speeches, and 85.5 percent fall within the seven families in the table above. The remaining speeches sit outside this comparison.

3.1Call it a threat and you get regulation; call it an opportunity and you get enablement

Join diagnosis to prescription and the path is close to automatic. Threat-diagnosing speeches lean strongly toward regulation and restriction, and opportunity-diagnosing speeches overwhelmingly favor enablement and investment. The paper is explicit that this comparison is descriptive and based on mutually exclusive primary frames. The conclusion the authors draw from it is the title of this section: whether AI is called a threat or an opportunity is not a preliminary to the political conflict but a substantial part of it.

This path is decisive for explaining the empty compensation box. Even the radical left, the most contestatory family, attaches regulation and restriction rather than compensation to its threat diagnosis. The paper's conclusion says that this family, more than any other, pairs a threat diagnosis with a regulation-and-restriction response rather than settling for compensation alone. Compensation is thin not only because the side pushing adoption is strong. The side trying to hold it back is not choosing compensation either. Greens lean the same way on diagnosis but split their response evenly between regulation and enablement, which places them between the radical left and the adoption-oriented mainstream.

3.2The radical right did not politicize this as a threat to workers

Existing research has repeatedly found that workers exposed to automation are more likely to support radical-right parties. At the party level, that link does not appear. Radical-right opportunity framing closely resembles that of the mainstream right, and threat is only modestly more common. The paper reads this as a matter of repertoire. National competition, sovereignty, and the danger of falling behind are languages the radical right already handles well, and they fit technological adoption; workplace regulation and institutionalized labor protection sit less easily on that repertoire.

The intermediate position of social democracy is explained through the same lens. It is the familiar approach to structural economic change: the broad transformation is accepted, politics manages the transition, and regulation is reserved for the most damaging consequences. The result is that the strongest parliamentary challenge rests with the radical left rather than with the party family that has historically represented organized labor. On this issue, in other words, the side that has spoken for workers is standing in the position of managing adoption.

The authors widen the radical-right case into a general account of how new issues form. Because AI's material consequences have not stabilized, parties are free to connect it to the conflicts and reputational resources they already hold. National modernization for the mainstream right, worker vulnerability for the radical left, managed transition for social democrats. The same technology, as they put it, is becoming several different political objects at once, depending on who is describing it.

3.3Being in government tilts a party toward adoption

One further gradient survives controls for party family, country, and year. Government-party speech leans more than opposition speech toward framing AI as an opportunity and emphasizing enablement and investment. Opposition parties sit closer to threat and regulation. Training barely responds to the distinction. The paper states twice that this comparison is associational and does not identify the causal effect of entering office. Reading it as consistent with the pressures of governing responsibility, such as national strategy, economic performance, procurement, and public-sector modernization, is as far as the authors go.

3.4What one speech actually looks like

Everything so far has been aggregate. To show what the individual speeches behind the aggregate look like, the paper picks five and lays them out in a table. What appears in the table is not verbatim quotation but content the authors paraphrased, and it should not be read as though it were in quotation marks.

Party Country and year Coded response Content as the paper paraphrases it
Die Linke Germany, 2025 Regulation and restriction Frames AI as a risk to workers and argues for public rules or constraints on deployment
Australian Labor Party Australia, 2026 Regulation and restriction Links AI to risks in employment or creative work and calls for tighter limits or safeguards
Liberal Party Canada, 2025 Enablement and investment Connects public AI investment to innovation capacity and the creation of future high-value jobs
Labour United Kingdom, 2025 Training Acknowledges AI displacement while emphasizing skills, transition support, and free AI training for workers
NDP Canada, 2025 Compensation Argues that AI can replace workers and links the response to employment insurance and income supports

The bottom row is what the 2.3 percent box looks like from the inside. Displacement is acknowledged, and the instruments that follow are employment insurance and income support, tools the welfare state has had for half a century. Look at the second row and creative work has arrived on the regulation side of the vocabulary. Those two rows preview both of section 4's findings: how thin the compensation box is, and the fact that the people who make the data reach the stage only through the regulation box.

Canada's Parliament Hill Centre Block — the parliament where both an enablement speech and a compensation speech were coded in the same year
▲ Canada's Parliament (Centre Block). Two of the five examples in the table above came from here. A Liberal Party enablement-and-investment speech and an NDP compensation speech sat in the same legislature in the same year, 2025. | Source: Wikimedia Commons
4

Open the regulation frame and copyright comes first

Of the four boxes, the one that looks closest to the question of what data workers are owed is regulation and restriction. The paper opens that 21.8 percent box and shows what is largest inside it. The values in this section have a different standing from those in the preceding sections. The authors embedded the English summaries retained during annotation, clustered them separately within each frame, and had humans name the clusters, and they state that these clusters are used only for interpretation and do not enter the statistical models or hypothesis tests. The evidence they offer for them stops at stability: rerunning across random seeds produced mean adjusted Rand indices between 0.889 and 0.928. The numbers below are best used to read direction, not to argue from a decimal point.

4.1The second-largest worry in diagnosis is last in prescription

What is largest inside the threat diagnosis? Talk of jobs disappearing and incomes becoming unstable takes more than half of threat content. Workplace control comes next, meaning concern about surveillance and about how algorithms handle people on the job. Inside the regulation response the ordering is different. The largest item is copyright, the creative sector, and media protection, and the box that answers workplace control, which covers algorithmic management, platform-work safeguards, and human oversight, is the smallest of the four.

Inside the threat diagnosis Denominator: threat-frame mentions Inside the regulation response Denominator: regulation-frame mentions Job displacement, income insecurity 53.1% Workplace control 19.5% · 2nd Falling behind in the AI economy 14.8% Skills gaps 12.5% Copyright, creative sector, media 36.1% · 1st Labor-market safeguards 27.2% Innovation-compatible regulation 23.7% Algorithmic management, oversight 13.0% · last 2nd ↔ last The two columns have different denominators

Redrawn from the values in section 4.5 and Figure 11 of the paper. The two columns take threat-frame mentions and regulation-frame mentions respectively as their denominators, so 19.5 percent and 13.0 percent cannot be subtracted from or divided by one another. All this figure says is that the rankings do not line up, and both columns come from the exploratory analysis rather than the validated coding. The last item on the right is the paper's "algorithmic management, platform-work safeguards, and human oversight."

It is the authors who point out this correspondence. The paper describes algorithmic management, platform-work safeguards, and human oversight as the topic closest to the surveillance and workplace-control concerns raised on the threat side, and then writes that it is the smallest category inside regulation. That is where the paper stops. Putting the two distributions side by side and saying that the second-largest worry in diagnosis finishes last of four in prescription is our sentence, and the paper does not name this a mismatch between diagnosis and prescription.

The opportunity diagnosis is worth opening briefly as well. If threat is one story, opportunity is two. Competitiveness and innovation take 37.8 percent of opportunity mentions and productivity, job creation, and labor-shortage relief take 34.9 percent, close to equal, with infrastructure and skills formation trailing well behind. Threat is largely one story about disruption to income and livelihoods, on the authors' summary, while opportunity is two, one about national competitiveness and one about firm-level productivity. These values come from the same exploratory analysis.

What the smallest box, algorithmic management, actually contains is well illustrated by an example from the paper's own introduction. Green and left parliamentarians in Europe pressed for stronger protections against algorithmic management at work, including rights to information, explanation, and human oversight over automated workplace decisions. The introduction sets an example in a very different register right beside it. Singapore has made AI adoption a centerpiece of its economic and workforce strategy, and parties around the world echo that language of investment, competitiveness, and modernization. Facing the same technology, one side wants conditions attached to deployment and the other wants adoption accelerated. The 55.2 against 21.8 measured in section 1 is how loud each of those two voices is.

4.2Copyright is the language of ownership, not of payment

That copyright leads the regulation box reads at first like good news, since it seems to mean that what creators and data providers are owed has made it onto the political stage. The paper describes this topic as a legacy concern that generative AI has revived rather than invented. And the question copyright answers is whose it is, not how much to hand back. The distance between those two questions is shown most sharply by the largest sum of copyright money ever produced, which was produced outside any parliament.

On 20 July 2026 the U.S. District Court for the Northern District of California granted final approval to the class settlement in Bartz v. Anthropic. The total is 1.5 billion dollars, covering 482,460 works, with a claims rate of 91.3 percent and an estimated gross distribution of roughly 3,100 dollars per claimed work. It is the largest copyright class settlement in U.S. history. Look at the nature of that money, though, and the direction changes. The June 2025 ruling at first instance held that using lawfully acquired books for training was fair use, as was digitizing purchased print copies, and declined to find fair use only for the path that ran through pirate sites into a central library. The 1.5 billion dollars is therefore compensation for the illegality of an acquisition route, not a price set on the labor of writing or the labor of making data.

Phillip Burton Federal Building and United States Courthouse — where the Bartz v. Anthropic settlement was granted final approval
▲ The Phillip Burton Federal Building and United States Courthouse in San Francisco, home of the U.S. District Court for the Northern District of California, which granted final approval to the $1.5 billion Bartz v. Anthropic settlement. | Source: Wikimedia Commons

Even the largest payment copyright has ever produced was not payment for creative or data labor. That politics speaks about copyright more than anything else inside the regulation box comes close to meaning that what the people who made the data are owed gets handled only where it can be reduced to ownership. Labeling, review, and feedback work, which does not reduce to ownership, does not fit into this vocabulary at all.

How much the language of copyright actually moves, and who holds that language, is something the Pebblous blog has examined in a piece on bargaining power in training-data licensing. The argument running the other way, that holding copyright tightly is not always to a creator's advantage, is in a piece on copyright market design and the penalty on originality.

4.3The compensation box is not only small but thin

Open the 2.3 percent compensation box and there are only two strands inside. Income support and safety nets take 64.0 percent of compensation mentions and reduced working time takes 36.0 percent, and the two topics together amount to 125 mentions in absolute terms. That is the whole of compensation talk left by 33 parliaments over three years and four months. The paper's assessment is level: what little compensation politics exists is also thin substantively, largely restating conventional welfare-state instruments rather than developing AI-specific remedies such as robot taxes or productivity-sharing schemes.

The explanation the paper offers for the small size is cautious. Stable constituencies of AI losers may not yet have formed at sufficient scale to organize conflict around post-disruption repair, so party competition is already developing further upstream, over how a still-unsettled transformation should proceed. On that reading, compensation is thin not because politics rejected it but because politics may not have reached it yet. The few compensation mentions that do exist are not spread evenly across parties but concentrated on the left, and even there they remain clearly secondary to regulation and restriction.

The paper has drawn a boundary of its own that cuts against that regret. The robot tax whose absence that sentence regrets is already classified as an example of the regulation and restriction frame in the same paper's coding scheme. Look again at the coding table reproduced in section 1 and the last item on the list of regulation examples is robot taxes. The definition in the body runs the same way, placing taxes intended to discipline labor-replacing adoption under regulation while compensation covers income protection, redistribution, reduced working time, and basic income.

Where the robot tax appears What it is treated as there
Table 1 coding scheme, regulation and restriction row One of the headline examples of the regulation frame
Narrative in the section 4.5 exploratory analysis An AI-specific remedy whose absence from the compensation box is regretted

If a legislator had raised a robot tax on the floor, the speech would have been coded as regulation, not compensation. This is less an authorial slip than what happens when a new instrument does not sit neatly in an older classification, and how much this boundary contributed to the apparent thinness of the compensation box cannot be checked without the appendix. What the debate around robot taxes has looked like so far is set out in the Pebblous blog piece on token taxes and the human share.

Outside parliament, matters look somewhat different. A policy document OpenAI published on 6 April 2026 proposed shifting the tax base away from payroll taxes toward capital income, corporate taxation, and new taxation of automated labor, and included a 32-hour week experiment without pay cuts and a citizen-equity public fund. That the most concrete AI-specific compensation instruments came from an AI company rather than from any political party overlaps exactly with the empty space this paper counted in legislatures.

4.4The 55.2 percent is not a subsidy to business

Read this far and it is easy to reach for a schema in which enablement is the language of the right and regulation the language of the left. The internal composition of the enablement box breaks that schema. Growth and national investment take 39.1 percent of enablement mentions and adoption support, productivity, and public-sector capacity take 31.3 percent, and the paper writes that these two topics together account for seven in ten enablement mentions. AI infrastructure and industrial strategy follow at 17.2 percent and SME adoption support at 12.4 percent.

The authors' reading is this. Enablement is directed at least as much toward building state and public-sector capacity as toward supporting private firms. Public authority is mobilized to expand the state's own technological reach, not only to clear the way for business adoption. That gives enablement a public-investment register available to the left as well as the right, and the paper adds that this may be why social democrats and greens, despite their threat-oriented diagnoses, still register substantial enablement in their response mix. Reading the size of enablement as evidence that parliaments have sided with industry misrepresents the paper.

4.5Training is large because both sides use it

The box not yet opened is training. At 20.6 percent it is nearly the size of regulation and restriction, and it does not discriminate between party families. It appears at broadly similar rates in every one of them. The paper locates the reason in this instrument's capacity to serve several projects at once. Building the capacity adoption requires is training; helping unsettled workers move to the next position is training; developing national skills and research capacity is training. The side pushing forward and the side holding back can use the same word in their own senses, so no fight breaks out here.

Open the box and those three strands are visible. Employee reskilling takes 45.6 percent of training mentions, AI talent and research capacity 28.9 percent, and schools and future skills 25.5 percent. These values also come from the exploratory analysis. Look for data labor here and you come away empty again. Training is a story about moving people to their next job, not a story about setting a price for the people currently making the data. That is the result of opening all four boxes. The vocabulary for payment is absent from the largest box, the second, and the third, and the fourth box is 2.3 percent.

5

Can these results be carried over to data labor?

That is as far as the paper goes. This report takes one step past it. What the paper investigated is the vocabulary parliaments use for AI's effects on work in general; what this report wants to read is whether what labelers, creators, and data providers are owed exists inside that vocabulary. The two questions overlap without being the same. How far the overlap runs is taken up here in three parts: what was studied, where the stage ends, and what two countries show.

5.1This paper did not count data labor

The unit of analysis is the speech, and what was coded is how that speech treats AI and work. Nowhere in the paper is there a passage counting labelers, content moderators, or reinforcement-learning feedback workers, or dealing with their wages. Antonio Casilli's Waiting for Robots: The Hired Hands of Automation, a book that addresses microwork and data labor head on, passes by once in a citation string in the introduction. That is itself a trace of the fact that this paper did not take that labor as its object.

So the fact that algorithmic management is the smallest topic inside regulation at 13.0 percent is not a count of labelers. What that box holds is the problem of algorithms managing people in the workplace, platform-work safeguards, and human oversight, and the labor of making data has no box of its own anywhere in this classification. Counting an absence requires a box to count it in, and this instrument has no such box.

The net and the box are two different things. The relevance definition in section 2 includes wages and worker protection, so had a legislator spoken about what labeling workers are paid, the speech would have been caught in the net. The problem starts after it is caught. The place it would sit is algorithmic management, or compensation, or copyright, and no name anywhere in this classification points at the labor of making data itself. The reason it goes uncounted is not that the net has wide holes but that there is no name to count it under.

The authors themselves pin down what this data cannot say, in their conclusion.

"While our evidence cannot establish a direct representation gap, since demand- and supply-side studies observe different populations, countries, and periods, it nonetheless leaves open the possibility that public unease about AI is outpacing the positions parties have so far built around it."

That sentence applies to this report as it stands. The thinness of the language of compensation in parliament does not yield the conclusion that people do not want compensation. What it yields is only that the language currently sits thinly on the supply side of the stage.

5.2The instrument measures lower-house plenary debate only

The boundary of the stage is drawn narrowly too. To maximize comparability, the authors narrowed the target to lower-house plenary debate, or the national unicameral chamber where applicable, as stated in section 3. In the limitations paragraph of the conclusion they then list for themselves what that choice misses.

"Parliamentary speech does not directly measure enacted policy, whose formation also depends on coalition bargaining, veto institutions, bureaucracies, courts, unions, employers, and organized interests."

Courts, unions, employers. Nearly every place where rules about paying for data are actually being written right now is on that list. Hold the apparent counterexamples up against this boundary one by one and most of them fall outside the instrument.

Case where payment rules are being written Where it is made Inside this instrument?
Bartz v. Anthropic settlement U.S. federal court Outside. The paper lists courts in its limitations paragraph
SAG-AFTRA collective agreements with AI clauses Collective bargaining Outside. Unions and employers on the same list
California AI surplus-sharing bill U.S. state legislature Outside. The sample is national lower-house plenary debate
EU AI Act and the platform work directive EU institutions Outside, although member-state floor speeches on domestic implementation are inside
Kenya AI and Emerging Technologies Policy Kenyan ministry Outside. A ministry document, not parliamentary speech
The Artificial Intelligence Bill, Kenya 2026 Kenyan Senate Outside. Not the lower-house plenary
South Korea's AI Framework Act National Assembly of Korea Outside. Korea is not among the 33 countries
News and music training-data licensing deals Private contract Outside. Not a political stage at all

How narrow the boundary is can be seen in the paper's own opening paragraph. The first case the introduction offers when announcing the arrival of this politics is a report by U.S. Senator Bernie Sanders and the AI data center moratorium bill Sanders introduced with Alexandria Ocasio-Cortez. The U.S. Senate is not in this paper's corpus. The protagonist of the first paragraph is not caught by the instrument in the body, which is not a flaw in the paper but a design choice the authors made explicitly in order to measure 33 countries with the same yardstick. What we do need to know is what else we are carrying along whenever we carry these results somewhere.

Each case in the table has been covered separately on the Pebblous blog. The surplus-sharing design a state legislature wrote is in the piece on the California bill, the consent and payment rules collective bargaining wrote are in the piece on the SAG-AFTRA agreement, and the hole in the EU's training-data disclosure duty is in the piece on the training-data summary gap.

5.3Kenya is inside the sample, and Kenya's attempt is outside it

The country where the meaning of this boundary shows up most clearly is Kenya. Kenya is on the list of 33. It is also the country where data annotation and content moderation work became the largest public controversy. And the document currently serving as the response to that controversy came from a ministry, not from parliament.

The Kenya Artificial Intelligence and Other Emerging Technologies Policy, issued in July 2026 by the State Department for ICT and the Digital Economy, finds that data annotation and content moderation workers face unique occupational risks, including exposure to harmful content, insecure working conditions, and limited labour protections, and that existing frameworks provide insufficient safeguards. Its prescription is a Fair Pay Reference Framework. The framework is to specify transparent pay benchmarks for data annotation, content moderation, and AI quality evaluation work, calibrated to international rates for equivalent roles, with a compliance reporting mechanism through which operators disclose their pay structures against those benchmarks. Statutory duty-of-care standards for the same workers appear alongside it, covering harmful content exposure, mental health support, fair contracting, remuneration, and workplace protections. The scope named in the document includes foreign entities whose AI systems are procured or deployed in Kenya, along with data annotation providers.

Nairobi skyline — capital of Kenya, where data annotation and content moderation labor became the largest public controversy
▲ Nairobi. Kenya sits inside this paper's 33-country sample and is also the site of the largest controversy over data annotation and content moderation labor. The pay rules for that work are being written in a ministry policy document, not on the floor of parliament. | Source: Wikimedia Commons

This document is a policy direction rather than a law, and the actual obligations will be created by the legislation and subsidiary rules that follow. It is also invisible to this paper's instrument, because it is a ministry document. On the parliamentary track, meanwhile, Kenya has The Artificial Intelligence Bill 2026, published as Senate Bill No. 4 on 19 February 2026 and given its first reading on 2 April. The bill digest issued by the Parliament of Kenya records the originality of academic and creative work and the disruption of employment opportunities as background, but the words worker, working conditions, and data annotation do not appear in that digest. Risk-based regulation with an AI Commissioner is the bill's backbone.

An attempt to actually write the payment for data labor into rules is under way in a country that is inside the sample, and that attempt sits in a ministry policy document while the parliamentary bill sits in the Senate and its digest contains no word for labor. How Kenyan legislators spoke about labelers on the floor cannot be learned from this paper, because it does not publish speech content by country. What we know reaches as far as this: Kenya is in the sample, a Kenyan ministry is writing wage rules for data labor, and those rules are being made outside this instrument.

Not knowing now does not mean not knowing later. In its data availability statement, the paper says the data and code needed to reproduce the analyses will be made publicly available in an open repository upon publication. At that point it will be possible to count directly which boxes the AI-work speeches of Kenyan legislators were coded into. This report stops here because of when it is being written, not because of what the material is.

5.4Korea is not in this sample

Finally, Korea. Korea is not among the 33 countries, so this paper's conclusions cannot be applied to the National Assembly as they stand, and neither is there any evidence in this data that Korea is heading in a different direction. It simply was not measured. What vocabulary has begun to be used domestically for high-impact AI is set out separately in the Pebblous report on high-impact governance under the AI Framework Act. Adjacent material worth reading alongside this paper includes a report on the mismatch between AI exposure and union protection and one on how the expert data labor market prices its work.

6

Why Pebblous is watching

Pebblous diagnoses data and issues quality reports on it. What 33 parliaments said touches that work at the level of timing rather than of regulatory forecasting. An organization that buys and makes data will not learn from this paper when payment rules will arrive or what shape they will take. What it can take away is that there is currently no basis for a schedule that assumes them.

6.1What is not priced does not get recorded

The absence of payment rules comes back as a quality problem in the end. Data whose contributors are not priced has no reason to record who put in what and how much, and data with no such record cannot later have a particular contribution erased or repriced. That copyright leads the regulation box reads differently from here. If ownership is the only vocabulary politics has for payment, then labeling, review, and feedback, which do not reduce to ownership, never even become objects of record. The holes in a data lineage are usually in that spot.

The Bartz v. Anthropic settlement is the most concrete evidence for this argument. What divided that case was not which book someone had written but by which route that book became training material. The record of the acquisition route is what decided whether compensation was owed, which shows the ordering: the record comes before the payment. How the fair-use judgment at the training stage and the gap at the acquisition stage come apart is treated separately in the Pebblous report on training-stage fair use and provenance records.

6.2When the rules do arrive, the first thing they ask for is records

Saying do not wait is not saying ignore the rules. Looking at what the rules ask for when they arrive makes the present task clearer, if anything. Kenya's policy document from the previous section is a good example. It does not contain a wage framework alone. It also commits to establishing minimum national standards for AI-relevant dataset documentation, provenance, quality assurance, representativeness, custodianship, and update cycles. That the clause requiring pay disclosure and the clause creating dataset provenance standards sit in the same document is not a coincidence. Deciding who gets paid how much requires first having a record of who put in what.

The items a data supply chain designer can act on now are therefore largely settled. How far the scope-of-use clause in the contract reaches across training, retraining, and derivative models; what identifier contributors are recorded under; how long provenance records are retained; whether there is a procedure for obtaining consent again when a model is retrained. None of the four requires waiting for a new institution, and all four can go into the next contract. The later the rules arrive, the wider the distance grows between organizations that have written them and organizations that have not.

6.3Contracts work in the space politics left empty

Reduce what this report counted to one sentence and it comes to this. The language for repaying labor was 2.3 percent of response-frame mentions across 33 parliaments, and what it contained was income support and reduced working time and nothing else. The labor of making data has no box of its own anywhere in this classification. What fills that empty space at present is court settlements, collective agreements, and private contracts. None of them is a stage designed to write payment rules fairly, but they are where the prices actually get set.

What an organization handling data can build now, then, is the record that will hold those contracts up. A state in which it can answer what data came in from where and how, who touched it, and under what terms it may be used. This is not the kind of thing to wait for an institution to settle, and it is also the first item an institution will ask for when it arrives. Whoever built records while there were no rules is the one who can carry the rules once they exist.

Please read the paper's findings and Pebblous's interpretation as separate things. Sections 1 through 4 report what Chueri and Törnberg measured, while sections 5 and 6, which carry it over into the question of what data labor is owed, do something the paper did not do. The figures from the paper in the body were checked directly against the full text of the version 1 preprint, and the contents of the Kenyan policy and bill were verified in the originals published by the government and the parliament. Sources for the remaining external facts are given in the body. Thank you for reading this far.

R

References

The figures in this report come from three streams. Values internal to the paper were carried over from the HTML version 1 and the 32-page PDF of the preprint, checked directly down to section, table, and figure number. The Kenyan policy document and bill digest were verified by downloading the original PDFs published by the government and the parliament, accessed on 6 September 2026. The remaining external facts were cross-checked across multiple outlets, and the nature of each check is recorded in the entry.

The backbone of this report

  • 1.Juliana Chueri (Vrije Universiteit Amsterdam), Petter Törnberg (University of Amsterdam). "Meeting the Coming Wave: The Emerging Politics of AI and Work across 33 Parliaments." arXiv: 2609.02296, v1 submitted 2 September 2026, cs.CY. Verified in this version: corpus size, the response-frame distribution in section 4.3, the exploratory topic shares in section 4.5, the coding scheme in Table 1 and the political question each frame answers, the sample stages in Table 2, the five illustrative claims in Table 3, the prior expectations and theoretical account in section 2, the relevance definition and party-family identification rates in section 3, the country profiles in section 4.1, the human validation figures, the limitations and conclusion passages quoted in section 5, and the data availability statement. The appendix is cited five times in the body but appears neither in the HTML version nor in the 32-page PDF. No peer-review status is stated.

Scholarship the paper is in dialogue with (verified in the paper's own reference list)

  • 2.Stefan Thewissen, David Rueda. "Automation and the Welfare State: Technological Change as a Determinant of Redistribution Preferences." Comparative Political Studies 52(2), 1670–1706, 2019. The starting point of compensation-centered political economy, and the position this paper defines itself against.
  • 3.Thomas Kurer, Silja Häusermann. "Automation Risk, Social Policy Preferences, and Political Participation." In M. R. Busemeyer (ed.), Digitalization and the Welfare State, Oxford University Press, 139–156, 2022.
  • 4.Marius R. Busemeyer, Mia Gandenberger, Carlo Knotz, Tobias Tober. "Preferred Policy Responses to Technological Change: Survey Evidence from OECD Countries." Socio-Economic Review 21(1), 593–615, 2023.
  • 5.Marius R. Busemeyer, Tobias Tober. "Dealing with Technological Change: Social Policy Preferences and Institutional Context." Comparative Political Studies 56(7), 968–999, 2023.
  • 6.Zhen Jie Im, Nonna Mayer, Bruno Palier, Jan Rovny. "The 'Losers of Automation': A Reservoir of Votes for the Radical Right?" Research & Politics 6(1), 2019. Background needed to read the reversal in section 3.2, where the radical right does not politicize AI as a threat to workers.
  • 7.Massimo Anelli, Italo Colantone, Piero Stanig. "Individual Vulnerability to Industrial Robot Adoption Increases Support for the Radical Right." PNAS 118(47), e2111611118, 2021.
  • 8.Aina Gallego, Thomas Kurer. "Automation, Digitalization, and Artificial Intelligence in the Workplace: Implications for Political Behavior." Annual Review of Political Science 25, 463–484, 2022.
  • 9.Lance Y. Hunter. "Political Ideology, Artificial Intelligence (AI), and Labor Markets: How Political Party Members Perceive AI's Effects in OECD Countries." Political Research Quarterly, 2026. The theoretical source of the diagnostic and response frame classification.
  • 10.Fabrizio Gilardi, Meysam Alizadeh, Maël Kubli. "ChatGPT Outperforms Crowd Workers for Text-Annotation Tasks." PNAS 120(30), e2305016120, 2023. And Petter Törnberg, "Best Practices for Text Annotation with Large Language Models," Sociologica 18(2), 67–85, 2024. The methodological grounds the paper cites for language-model annotation.
  • 11.Antonio A. Casilli. Waiting for Robots: The Hired Hands of Automation. University of Chicago Press, 2025. A book that addresses microwork and data labor head on, and one that appears in this paper only once, in a citation string in the introduction.

Primary policy and litigation documents (accessed 6 September 2026)

  • 12.State Department for ICT and the Digital Economy, Republic of Kenya. "Kenya Artificial Intelligence and Other Emerging Technologies Policy," July 2026. PDF published at ict.go.ke. Verified in this original: the diagnosis of occupational risks facing data annotation and content moderation workers, the Fair Pay Reference Framework and its pay-disclosure mechanism, the statutory duty-of-care standards, and the standards for dataset documentation, provenance, and quality assurance. The file name begins with "draft," marking a version issued for consultation. Hourly-rate gaps between Kenya and the United States reported by several outlets could not be verified in this original and are therefore not used in the body.
  • 13.Parliament of Kenya, The Senate. "Bill Digest: The Artificial Intelligence Bill, Senate Bills No. 4 of 2026." PDF published at parliament.go.ke. Sponsor Sen. Karen Nyamu, published 19 February 2026, first reading 2 April 2026. The absence of the words worker, working conditions, and data annotation from this digest was also verified in the original.
  • 14.Final approval of the class settlement in Bartz v. Anthropic. U.S. District Court for the Northern District of California, 20 July 2026. Total 1.5 billion dollars, 482,460 works covered, claims rate 91.3 percent, estimated gross distribution of roughly 3,100 dollars per work. Cross-checked against Authors Guild, Authors Alliance, and JURIST. The scope of the June 2025 fair-use ruling at first instance was verified in the same sources.
  • 15.OpenAI. "Industrial Policy for the Intelligence Age," 6 April 2026. Contains the proposal to shift the tax base from payroll taxes toward capital income, corporate taxation, and taxation of automated labor, along with a 32-hour week experiment and a citizen-equity public fund. The official page refuses automated retrieval, so the original could not be checked directly; this entry rests on summaries in multiple outlets.
  • 16.U.S. Senate Health, Education, Labor and Pensions Committee Minority Staff. "The Big Tech Oligarchs' War Against Workers," October 2025. The report that opens the introduction of reference 1.

Adjacent Pebblous writing