Executive Summary

Anthropic released a robot exposure index at the end of September 2026, and the figure that travelled was 74%: the share of US physical tasks that today's robots could perform. Coverage attached a second sentence to it. Even in the fastest scenario, the study said, half of all physical work does not become cheaper than human labour until 2050. That became a story about reassurance, about blue-collar workers having decades of room. The same study also wrote down a date for its baseline scenario. That date is 2085, and it appears nowhere in the coverage.

Traced backwards, the chain takes in only four kinds of data from outside: job descriptions, employment and compensation statistics, population survey data, and a robot price series. Every one of the ten steps between them is an estimate produced by Claude Opus 5. The evidence on the robot side is not field logs either; it is product pages and trade press. In the public dataset, close to one citation in three points at the website of the company that built the robot. The work entered through text written by whoever drafted the job description. The robots entered through text written by the firms that sell them. Both sides are documents. And a paper that ran the same kind of rubric with the grader model swapped out had already appeared five months earlier. When the grader changed, close to half of the task-level ratings changed with it.

Everything above is fact. Here is the interpretation. The authors wrote down their own limits. Without the ratings that lean on a robot doing some similar task, the appendix says, the exposed share falls from about three quarters to about one half. The sentence saying that much physical work is doable but not automatable at scale with today's capabilities is theirs as well. So is the caveat that the fifty-year backtest validated factory interiors, while the warehouses and roads this study is actually new about are the part history says little about. This is a study that drew its own sensitivity bands and its own range. What this report questions is not the authors' diligence but the distribution path that dropped the appendix and carried the headline. And underneath that, a more basic condition: at the place where the future of physical AI is being measured, there is no number that was counted in the field.

3/4 → 1/2

When transfer-based ratings are removed

A sensitivity the authors computed themselves. 24 percentage points hang on one judgment call

56.9%

Agreement when the grader model changes

From a study that ran the same rubric on other models. Three runs of one model agreed 99.0% of the time

80 / 105

Unstructured-environment tasks that are not driving

105 tasks were rated as doable in settings like public roads, and 80 of them have nothing to do with driving

30.8%

Evidence from robot makers' own websites

A lower bound, from sorting all 56,933 published citations by domain

1

Four Numbers, and the One That Went Missing

Anthropic published "What work can robots do?" on 30 September 2026. Russell Legate-Yang and Maxim Massenkoff wrote it: 45 pages of main text, 57 pages of appendices. This is a company research report, not a journal article. There is no arXiv number and no DOI, and the citation format the authors supply is the one used for online documents. What they did publish is the full set of prompts and the task-level dataset, and every recomputation in this report comes out of that release.

Yahoo Finance ran a story the following day, and that story carried three headline numbers. Robots can do 74% of physical tasks, those tasks account for 34% of working hours, and robots are cheaper than people on 0.3% of them. There are four headline numbers rather than three, though, and they do not share a denominator. Without that separation, the very next sentence goes wrong.

1.1Different Denominators Behind Each Number

The table below sets out the four denominators. 74% is a share of US physical tasks. 34% and 0.3% are shares of all working hours, physical or not. Because those last two do share a denominator, they can be read side by side, and the report's own sentence makes exactly that contrast. Work that robots could in principle do amounts to 34% of all working hours, while the work on which a robot is cheaper than a person today comes to 0.3% of the same total.

Headline What it counts Denominator
74%Tasks rated as performable by today's robotsUS physical tasks
34%Working hours those tasks absorbAll working hours
0.3%Tasks where a robot costs less than a personAll working hours
81%Tasks exposed to robots or to large language modelsAll task hours

Taken from the report's main text and its key findings list. The key findings give the fourth number as "about 80%" while page 14 gives 81%. Same value, rounded and unrounded; this report uses 81% throughout.

Pebblous has looked at an index built on the same occupational database before. The occupational map of agentic delegation covers text work. This report pulls the physical tasks out of the same database and attaches robots to them instead. Anyone reading the two indices together should start by noting that they are scored over different objects.

1.2The Axis Is How Much the Workplace Must Be Changed

The most important device in the study's design is a four-tier scale. The press summarised it as "four levels of environmental control," which is accurate as far as it goes but easy to read with the axis inverted. The question the scale asks is not how controlled an environment already is. It asks how much a person has to modify the environment before a robot can do the work there. The tier that needs the most modification is E1; the tier that needs none is E3. Exposure is therefore highest at E3.

Tier Meaning Of physical tasks Of all tasks
E0No robot can do it~25%12%
E1Only in a purpose-built robot setting (factory assembly line)50%23%
E2In a structured human facility (logistics warehouse)22%10%
E3In an unstructured environment (city streets)2%1%

E1 through E3 sum to 74% of physical tasks and 34% of all tasks. The E0 share lands anywhere from 24% to 30% depending on the weighting and base year, so only the employment-weighted figure is used here.

This is where the internal composition of the 74% becomes visible. Fifty of those percentage points are E1. The tier that requires building a new workplace around the robot takes up more than two thirds of the headline. The tier that matches what people picture when they hear "robots can do our jobs" — work done in an environment nobody rearranged, on a city street — is 2% of physical tasks.

Four Environment-Control Tiers (E0–E3) 25% 50% 22% 2% E0 E1 E2 E3 Can't do Purpose-built factory Structured facility Unstructured environment Retrofit need — High → Low
▲ Pebblous original diagram (environment-control tiers reconstructed) — of the 74%, 50 points sit in E1, purpose-built-only

1.3The Missing Number Is the Baseline Date

The report produced its arrival dates under two assumptions about how fast prices fall. In the baseline scenario, robot deployment costs decline 3% a year. In the rapid scenario, hardware prices fall 12% a year while the rate at which robots acquire new capabilities doubles. The year in which half of all physical tasks become cheaper than human labour is 2085 under the baseline and 2050 under the rapid scenario.

Coverage carried the second of those. The Yahoo Finance piece ran 2050 and "40 years" and did not print 2085. The E0 through E3 tiers are absent from it too, as is the sensitivity the appendix attaches to the 74%, as is the finding that regulation blocks 14% of physical tasks. The frame was reassurance: blue-collar workers have decades. That is not what this report disputes. The point is narrower. The 2050 doing the reassuring is the value from the scenario the report itself labels fastest, and the value the report treats as its default sits 35 years further out.

Carry only the headline and two things fall away together. Inside the 74%, E1 stops being distinguishable from E3; and between the two arrival dates, the baseline disappears. The rest of this report follows why those two matter, using the study's own appendices.

2

Where the 74% Came From

The 74% was not obtained by observing robots. For each task, Claude Opus 5 decided whether a robot available today could do that work, and those decisions were weighted by employment and summed. There is therefore only one way to read the number: look at what went into the decision, and look at the rule the decision followed. Both are public.

2.1Four Kinds of Outside Data

Work through appendices A to F, keep only the inputs the model did not produce, and four remain. O*NET 29.3, the US Department of Labor's occupational database. The Bureau of Labor Statistics figures on occupational employment, wages and total compensation. The Census Bureau's American Community Survey. And a robot price series going back to 1990. All 923 occupations and all 18,796 task statements come from the first of those.

Every step connecting those four inputs is a model estimate. The table below lists the ten of them. None of this is hidden; the authors attached the prompt for each step to the appendix, which is why an outsider can count them this way at all.

# What the model estimated Source
1Splitting tasks into physical, cognitive and interpersonal; assigning time and skill levelAppendix A.1
2Enumerating representative instances per task and weighting them by frequencyAppendix A.3
3Rating each instance's exposure tier, searching the web for robots and citing sourcesAppendix A.3
4The time share each task takes within an occupationAppendix A.1
5Annual output a person produces on that taskAppendix D.1
6Annual all-in robot cost per taskAppendix D.1
7Aggregation to the occupation, with de-duplication and a coordination surchargeAppendix D.1
8Classifying four kinds of barrier to automationAppendix C.2
9Describing ten clusters of work robots cannot doAppendix C.1
10Historical exposure tiers as of 1977, 1991, 2002 and 2008Appendix B.2

Compiled from appendices A through F. Outside data touches the chain in three places only: the task statements at step 1, the labour costs at steps 6 and 7, and the price trend.

Two Kinds of Documents Meet as a Judgment O*NET Task Descriptions 923 occupations · 18,796 tasks Robot Product Pages & Trade Press 56,933 cited sources Claude Opus 5 Judgment ~650K searches · 275K+ citations 74% of physical tasks Field-Measured Data no input this slot is empty
▲ Pebblous original diagram (estimation chain reconstructed) — both jobs and robots entered as documents; no field-measured value was input

2.2Evidence Required by Design

The model was not left to decide on its own. Appendix A.3 states the rule plainly.

"Claude is required to name a specific robot and provide direct source quotes to justify its claims."

The requirement actually ran. Counting the historical tiers as well, the pipeline consumed roughly 650,000 web searches, more than 275,000 citations and over 90,000 unique web pages. The present-day file alone, as published, contains 56,933 citations. Each rating carries its sources, and each source is tagged with a maturity level: deployment, commercial availability claim, or demonstration. This is a design that demanded evidence from the model and then left that evidence in a form outsiders can count. Without saying so first, everything that follows reads as an indictment, which it is not.

2.3"Whether the Workplace Would in Fact Be Rebuilt Is Not Part of the Judgment"

So what was the model actually rating? The prompt in Appendix F names two rules. The first is a cap: an instance cannot be rated above the tier of the environment it genuinely occurs in, so work that happens in a hotel corridor tops out at E2. The second rule is the one this report turns on.

"COUNTERFACTUAL. Below the cap, judge the instance as if it arose at the tier in question: … at E1 the instance starts with the work fixtured and presented as a purpose-built setting would present it. Whether the workplace would in fact be rebuilt is not part of the judgment."

That single paragraph fixes what kind of number the 74% is. E1, which holds half of all physical tasks, does not mean "a robot is doing this work in that workplace now." It means "if the work were fixtured and presented to the machine, a robot could finish it." Whether anyone would rebuild that workplace was explicitly excluded from the judgment. This is design rather than inference, and the authors did not bury it; they published the prompt.

A sentence the same authors put in Appendix C.2 is the reverse side of the rule. At least 42 percentage points of physical tasks are work that robots "can do in some environment but cannot automate at scale with today's capabilities." The distance between a rating of can do and a rating of worth doing is something the authors quantified themselves. Calling the 74% a measurement is therefore wrong. The accurate word is rating.

2.4What Kind of Writing That Evidence Is

The evidence also sorts by who wrote it. The 56,933 citations in the public dataset resolve to 11,174 unique domains, but at least 17,551 of those citations, 30.8% of the total, point at the domain of the company that built the robot in question. Vendor announcements routed through press-release wires and first-party domains buried in the long tail are not in that count, so it is a lower bound. The single most-cited domain is the trade publication The Robot Report, with 1,629 citations, or 2.86%.

One common summary needs correcting here. "It was measured from job descriptions rather than from field sensor logs" is half right. The work did come in through job descriptions, but the robots came in through product pages, trade articles and press releases, not job descriptions. Both sides are documents. The work entered the index as text written by O*NET's drafters; the robots entered as text written by their manufacturers and by trade media. Neither side was counted in a workplace.

Pebblous has retraced this configuration before, where Anthropic built an index about the economy using its own model. Our test of the company's AI usage statistics against an independent corpus is the earlier case. The difference this time is that the prompts and the task-level data were released together, which is what makes the recomputation in the next section possible.

3

The Authors Cut Their Own Number Twice

Every place where this report is easy to attack is already written into its appendices. The authors calculated two sensitivities that shake the headline and published both, and they drew the range of their own backtest. So there is nothing to expose below. What follows sets down what the authors wrote, what an outsider recomputed from the public data, and what never got written at all.

3.1Sensitivities the Report Computed on Itself

The first cut is in Appendix A.4. Where the evidence for a rating was not a robot doing that task but a robot doing a similar one, striking it out drops the exposed share of physical work "from roughly three quarters to roughly one half." Twenty-four percentage points therefore rest on a single judgment: whether this robot should count as able to do that task too. Striking out the ratings grounded only in demonstrations instead moves the figure by one percentage point, with the occupational rank correlation holding at 0.99. The gap in size between those two sensitivities is what identifies the weak link in the index.

Pebblous confirmed the same direction independently from the public dataset. Restricted to the instances that determined each tier, with the transfer-only ones stripped out, the exposed share of physical tasks drops from 66.5% to 49.3%, a fall of 17.2 percentage points. The two numbers should not be mixed. The authors' 24 points are employment-weighted and appear to come from re-rating the transfer-based judgments outright, whereas the Pebblous figure of 17.2 points is unweighted and comes from subtracting after the fact while leaving the published instance structure intact. Different methods, different magnitudes. The direction is what they share. The demonstration exclusion behaves the same way in both: the Pebblous recomputation puts that fall at 0.21 percentage points.

The second cut is in Appendix C.2. Physical tasks no robot can do in any environment come to 26%; physical tasks blocked from automation by capability come to 68%. The difference leaves at least 42 percentage points in between, and the authors' own gloss on that band is work robots "can do in some environment but cannot automate at scale with today's capabilities." Those 42 points sit inside the 74% headline.

Those 42 percentage points are not, however, the same kind of rating as the 74%. The barrier classification is a second, separate pass, and it asks a different question. For work robots do not do today, it first asks whether such a robot could be built with current technology, and where the answer is yes, it looks to comparable robots to judge what would get in the way. The authors flag that difference on the spot: the estimate is more forward-looking than the exposure measure, at the cost of being less grounded in today's robots. So when the two sensitivities are set side by side, the 24 points and the 42 points do not carry equal weight.

The maturity of the evidence behind the ratings is public too. Of the 56,933 citations, 36.7% are actual deployments, 45.2% are commercial availability claims and 18.1% are demonstrations. That the largest block is "we sell this" is not in itself a defect. That "installed" and "for sale" are different classes of evidence is something section 4 comes back to.

3.2The Inequality Sign, Pinned Down from Outside

A task's tier is set by weighted majority vote over its instance tiers. Appendix A.3 writes that formula, and both the prose and the equation say "at least half by weight." The Appendix F prompt that actually produced the data states the same rule as "more than 50 out of 100 points." At least versus more than: one inequality sign apart.

Which one the data followed can simply be counted. Implemented as "at least half" over the 7,594 published physical tasks, the formula puts 48 of them, 0.6%, at odds with the published tier — and all 48 are ties, with weights summing to exactly 50.0. Switch the sign to "more than" and all 7,594 match. The data followed the prompt; it is the methods description that is out of line.

The value of this finding is its kind, not its size. 0.6% does not move the headline. What matters is that releasing the full prompts alongside task-level data let someone outside identify the inequality sign the authors actually used. It also confirmed that a single sign accounts for all of it, with no trace of manual adjustment. Read this passage as evidence of reproducibility.

3.3A Dividing Line Nowhere in the Published Fields

How does the study handle robots under human control? The main text says a robot must operate "largely" without human control, and the performance condition in Appendix F is more precise: intermittent teleassistance is allowed, continuous teleoperation and step-by-step scripted control are not. The line falls between intermittent and continuous.

Checking that line in the public dataset runs into a wall. The autonomy flag on each source takes three values — autonomous, teleoperated, unclear — and none of them separates intermittent from continuous. The quotations are truncated at five leading words, so 56,758 of them, 99.7%, run to six tokens. Meanwhile 23.6% of exposed tasks have at least one source marked teleoperated or autonomy-unclear among the evidence that determined their tier. Weighted by time the figure is 23.8%, essentially the same.

That number does not mean the authors violated their own standard. The standard permits intermittent assistance, so a teleassisted source in the evidence is not by itself a breach. One thing does remain. Which side of the standard a given rating falls on cannot be determined from outside.

The prompts carry one more sentence. All three prompts in Appendix F contain the same sentence: "your assessment is used in economic research and audited by human reviewers." Sweep appendices A through F end to end, though, and no term for inter-rater agreement appears even once. No kappa, no repeat runs, no temperature or seed settings. The validation that does exist, in Appendix A.2, compares the classification axis — the split into physical, cognitive and interpersonal — against the BLS Occupational Requirements Survey, with correlations of 0.76, 0.89 and 0.78. Those are good numbers. They are not agreement on the exposure tier itself. The audit was promised to the model and its results were never reported to the reader.

3.4A Swap-the-Grader Experiment, Five Months Earlier

"What happens if you run the same prompt on a different model" is usually a question without an answer. This time there is one. In April 2026, NBER Working Paper 35110 took the same tasks from the same occupational database, used the same rubric wording, and changed nothing but the model. Its subject is language-model exposure rather than robots, but the method — have a model read task statements and assign tiers — is the same genus as this report.

What was measured Value
Cross-model agreement on task-level ratings56.9–73% (kappa 0.36–0.56)
Cross-model spread in mean occupational exposure3.6× (0.14 to 0.51)
Direction of disagreement2,728 tasks one model called exposed were unexposed to another; only 56 went the other way
Downstream regressions using those labelsIndividual-level coefficients differ by 2.4×, and at the regional level the sign flips
Three runs of the same model90.8–99.0% (kappa 0.85–0.98)

Yin, Vu and Persico (2026), NBER Working Paper 35110. The last row comes from Table A2 of the same paper: three repeated runs at temperature 0.

That last row is the decisive one. What breaks is the model, not the repetition. Run one model three times and the ratings come back almost unchanged; change the model and close to half of them move. The paper also reports that the disagreement does not cancel out in both directions, meaning one model grades systematically more generously. Anthropic measured neither quantity. Even if a repeat-run stability figure arrives later, that is the dimension this literature has already found not to be the problem.

Versions have to be stated alongside this. That paper used Claude 4.5; the report uses Claude Opus 5. They are not the same model, so nobody can assert that Anthropic's tiers wobble to that degree. What can be stated as it stands is that the grader scoring exposure highest in that comparison was a Claude model.

Closer evidence sits inside the report itself. Its language-model exposure measure comes from Eloundou et al. (2024), whose public file carries a human-rater column and a model-rater column side by side, and Anthropic chose the model column. Footnote 46 of the main text quotes the difference directly: computer programmers rank 25th by human raters and 6th by model raters, and poets and lyricists rank 11th and 14th. Recount all 923 occupations in that public file, as Pebblous did, and the footnote's examples turn out not to be exceptions. The model rater scored mean occupational exposure 8.6 percentage points higher than the humans, and 310 occupations — a third of the total — shift rank by more than 100 places.

The Eloundou paper itself carries one more thing. It published human-model agreement rates in a table, running from 65.6% to 82.1%. It reported that changing only the rubric wording on the same model flipped 24% of the labels. It stated explicitly that "we present the results from human annotators as our primary results." Its rubric even carried a numeric threshold: reduce the time to complete the task by at least half. Numbers are what went missing in the lineage.

3.5One Column the Whole Lineage Filled Is Empty Here

Occupational exposure indices have a lineage more than a decade deep, and that lineage has always put a number on how the labels were made trustworthy. The table below is that column.

Study Who made the labels Rater validation reported
Frey & Osborne (2013/2017)70 occupations hand-labelled by researchers, then a classifier100 split-half cross-validations, AUC 0.894
Webb (2020)No humans; deterministic overlap of patents and task statementsNo raters, so checked against historical outcomes
Eloundou et al. (2023/2024)Humans plus GPT-4Human-model and model-model agreement, 65.6–91.1%
Colombo et al. (2025, peer-reviewed)Three open-weight 7B-class modelsThree-model consensus published; three humans agreeing 71–75%
Schaal (2025)Three modelsCross-model correlation 0.60–0.81
Yin et al. (2026)Three modelsAgreement, kappa, repeat runs, downstream regressions
This report (2026)Claude Opus 5 aloneNo agreement, stability or cross-validation for the exposure tier

The Appendix A.2 validation in the last row compares the task classification axis against the Occupational Requirements Survey; reproducibility of the exposure tier itself is not reported.

The Colombo row is the one that catches. A far less well-resourced academic team used three open-weight 7B-class models and did the whole set — model redundancy, conservative aggregation, published consensus rates, validation by three humans — and gave as its reason that it wanted to guarantee the reproducibility of the research. Using several models, in other words, is not something only a large company can afford.

The price of the empty column is also on record. Ludwig, Mullainathan and Rambachan, cited in Anthropic's own Appendix footnote 10, wrote the answer down. Their paper splits language-model use in two. For prediction problems the condition is the absence of training leakage; for estimation problems, where the measurement of an economic concept is automated and fed into downstream analysis, the condition is a validation sample. Anthropic cited the first and did the second. The paper's simulations report that feeding model labels into a regression without correction badly degrades the actual coverage of a nominal 95% confidence interval, and that relabelling just 5% of the labels by hand restores it.

Five per cent of 7,594 physical tasks is about 380. Had the human-reviewer audit the prompts promised three times actually been run on 380 of them and the agreement rate written down, the standing of this number would be different. One step further would be too far, though. The simulation setup and this report's design are not the same, so the usable sentence stops at "there is no guarantee that a nominal 95% is an actual 95%."

3.6"Can Do" Has Conditions and No Numbers

There is a standard for performance. The Appendix F prompt defines "completing the instance" through three named conditions: finish the whole instance rather than a sub-step, from the start through to an accepted output; produce quality an employer would accept in place of that worker's output; and operate autonomously. The conditions are set out clearly in words. Numbers are what is absent. At what success rate, over how many attempts, with how many human interventions permitted, under what disturbances, at how many times slower than a person. Five blanks.

Nor does academia always fill those blanks with numbers. The leading manipulation benchmarks in robotics define success without fixing a trial count. ISO's robotics vocabulary standard notes that quantified autonomy metrics exist only on the medical electrical equipment side, and SAE's driving automation levels state three separate times that they do not provide specifications, being a taxonomy of design intent. The one precedent found with a number in it is a 2007 NIST document, which defined an intermediate autonomy level as "approximately 50% dependence on operator input." Anthropic's "intermittent teleassistance is allowed" leaves empty in 2026 a blank that was filled with a number in 2007.

One of those blanks collides with the report's own citations. The performance conditions include delivering the work "in a timeframe appropriate to the instance." Yet the Epoch AI analysis this report cites three times finds robots three to ten times slower than people, and footnote 43 of the main text says in its own words that a cleaning robot is "up to 10× slower than a human." Speed was written into the standard, and falling short on speed never cost a rating. The evidence requirement is asymmetric by design as well. Appendix F.3 demands at least one cited piece of evidence before a task may be rated as doable, and demands none for a rating of not doable.

4

Validated Inside the Factory, New in the Warehouse

One piece of this report is empirical work with nothing to criticise: Appendix B, which takes fifty years of historical data and tests what actually happened to occupations that scored high on exposure. This section introduces that work properly first and then turns to the range the authors drew around it. That the line around it was drawn by the report itself is what matters here.

4.1The Fifty-Year Backtest Is Real Empirical Work

The authors ran the same tier rating over job descriptions from five points in time — 1977, 1991, 2002, 2008 and 2026 — to build historical exposure, then tracked roughly 300 consistently defined occupations across three twenty-year windows. The results are clear. An occupation whose tasks were all exposed had wages 7.1% lower twenty years later (95% CI 5.4–8.9%) and employment down 34.2% (CI 16.4–52.0%). A placebo test passed: 1977 exposure has no effect on 1970s wages. The effect survives when a patent-based exposure measure is included alongside it, and when skill, offshoring and unionisation are controlled for.

One case carries more conviction than the numbers do. Machine operators, in the top 5% of exposure as of 1991, fell from roughly 100,000 in 1990 to roughly 50,000 in 2012 within the motor vehicle sector. Industrial machinery mechanics, at about a quarter of that exposure in the same industry, grew employment by 50%, and first-line production supervisors, barely exposed at all, grew 10%. This is evidence that the index does not point at just anything.

4.2Limits Drawn Around That Evidence by the Authors

The final paragraph of Appendix B.3 is the most honest sentence in the report.

"A caveat is that robots in this era were mostly used in controlled, E1 factory settings, so historical data say little about the E2 or E3 exposure that more adaptable, AI-controlled robots may create."

What was validated is the factory interior; what this study is new about is warehouses and roads. That range shows up inside the numbers this study produced as well. The headline coefficient of 34.2% is an equally weighted average of three windows, and window by window the employment effect shrinks from 50% in 1980–2000 to around 25% in the later ones. The reason the authors attach is decisive: the exposed set widened to include "new kinds of jobs, such as service work." In 1977, half of exposed occupations spent 95% or more of their working hours on physical tasks; by 2002 only 24% of them did. The further the exposed set moved from factory-ness, the smaller the coefficient got, and E2 and E3 are the set that moved further in precisely that direction.

That gap is older than this report. Acemoglu and Restrepo's 2020 study, the canonical empirical work on robots and employment, relies on International Federation of Robotics counts that explicitly exclude single-purpose machines — and the excluded example happens to be "automated storage and retrieval systems in warehouses." The logistics warehouse Anthropic offers as its archetype of E2 has never been counted in the historical data infrastructure at all. The effect that study reported is a 0.2 percentage point drop in the employment-to-population ratio and a 0.42% drop in wages per additional robot per thousand workers. The one-line caveat in Appendix B.3 is a limit of this whole field's data infrastructure, not of this report alone.

4.3Opening All 105 Unstructured-Environment Tasks

How big the genuinely new tier is can be counted directly. The public dataset rates 105 tasks as E3. That is 1.4% of the 7,594 physical tasks and 0.56% of all tasks. The main text summarises this tier as "mostly driving," which by task count it is not: 80 of the 105, or 76%, have nothing to do with driving. Bricklaying, wall finishing, operating paving and compacting equipment, collecting agricultural field data, inspecting construction sites, forest mapping — outdoor field work of every sort is mixed in. Weighted by time share, though, taxi, truck and shuttle driving take the top four slots, lifting driving to 39.7%. So the summary is half right.

The companies holding this tier up are few. Across 973 unique sources cited for E3 tasks there are 803 unique system names and 535 developers, which looks diverse, but weight by citation frequency and Waymo alone appears in 31 of the 105 tasks. And even in the tier with the highest exposure, 6 of the 105 have no evidence that could be called an actual deployment, resting entirely on commercial availability claims or demonstrations.

4.4A Janitor's Day: Where the Method Holds and Where It Runs Out

To bring the methodology down to somebody's actual work, pick one occupation and open it task by task. Janitors and cleaners make a good case. Three tasks that take up a large share of working hours received three different tiers, and the public dataset names the specific robot cited behind each rating.

Task Hours Tier Product actually sold Model's cost verdict
Floor cleaning27.3%E2Yes (Tennant, Avidbots, SoftBank, Gausium)40% more than a person
Rubbish collection15.2%E1None15× a person
Restroom cleaning16.7%E0Three start-ups citedCannot do it

Machine names read out of the source fields for those tasks in the public dataset. This is name-matching rather than statistical recomputation.

In the first row the method is accurate. Floor-cleaning robots really are sold. Convert the report's "40% more than a person" using median total compensation for janitors and that task's hour share and it comes to roughly $16,800 a year, which lands in the middle of the $13,500–$21,900 band for actual total cost of ownership on commercial floor-cleaning robots. Rent one on subscription and it comes in below the model. In the one task where field prices allow a check, this study's cost estimate holds.

The second row is different. Open all nine sources cited for "gather and empty trash" and not one of them is a robot that walks a building emptying waste bins. They are a waste-sorting robotic arm on a recycling-facility conveyor, a bag inserter on a packaging line, a mobile base that moves waste containers, a hospital supply delivery cart, a restroom scrubber, an outdoor litter-picking machine for lawns, and two demonstrations. The "15× a person" was therefore not measured on an expensive rubbish robot. No such robot exists, so the price of the nearest available machine went in instead.

That is still not an error. The authors rated the task E1, and a recycling facility is exactly an E1 environment, so the tier matches its evidence. They also published the dataset, so anyone can check those nine sources. What happened is this. The model answered honestly that robots doing this work exist only inside dedicated facilities and cost fifteen times as much, and that answer became "robots can do a janitor's job" on its way into a headline.

The third row is evidence in the opposite direction. Restroom cleaning cited three start-ups claiming commercial availability and still came out at E0, not doable. Vendor material did not automatically push a task into the exposed tier, and the weighted majority vote behaved conservatively there. Taken together the three rows reveal the character of the method. Where a product is in volume production it matches field prices, and where no product exists it substitutes the price of the nearest industrial installation. The trouble lies not in the method but in the fact that those two situations are indistinguishable inside the 74% headline. And the authors had already supplied the means to tell them apart, in the tiers and in the public dataset.

4.5Robots Named and Robots Actually Installed

The examples the main text names can be set against installation records. Amazon's shelf-stocking robot is in two sites, and the company's own verb is "plans to deploy." The figure of a million Amazon robots traces not to Amazon but to a news report; the definition of robot it counts has never been published, and it has gone fifteen months without an update. The Boston Dynamics warehouse robot is cited to a product page, and the only confirmed installation quantity on public record is 22 units at one European retailer. The humanoid offered as a car-factory case was a single unit — both companies write it in the singular — doing one pick-and-place task for ten months.

Lined up by scale it reads like this. Official IFR statistics put 2025 industrial robot installations at 603,307 units, and full-size humanoid sales in the same statistics at 7,000. Eighty-six to one. Those 7,000 were split among 174 manufacturers, about 40 apiece. The same federation wrote that humanoids are still "at the threshold of moving from the lab into pilots." The audited numbers are smaller still. Agility Robotics reported FY2025 total revenue of $1.78 million in its securities filings, and in Tesla's quarterly installed-capacity table the annual capacity cell for Optimus is a dash.

A structural property sits under all of this. Launch announcements stay searchable forever; quiet shutdowns do not. A robotic work cell Amazon unveiled as a flagship was marked four months later with an editor's note saying it is no longer in use. Any method that gathers robot capability by web search carries that asymmetry structurally. Whether that particular case actually entered this dataset could not be established, since the collection date relative to the note is unknown, so the claim here stops at the structure.

None of these installation figures licenses the conclusion that the 74% is wrong. Fifty of those percentage points are a tier that presupposes a purpose-built factory environment, and the tier for road-like settings is 2% of physical tasks. Moreover the study's paired figure — cost-competitive work at 0.3%, reaching 10% in roughly 40 years — matches those installation records exactly. What fails to line up is not the study but the summary that carried only the headline.

5

Between 2085 and 2050

Both dates came out of the same study. What separates them is not a worldview but two numbers: how fast prices fall, and how fast robots acquire work they could not do before. This section follows how far the grounding for those two numbers can be verified.

5.1Price Decline and Capability Growth Moving Together

2085 is not the answer to "what if prices fall 3% a year." The relevant paragraph of the main text lays capability growth down first and adds price decline on top. In 1977, robots could not do 62% of physical tasks; today's robots do most of that former set, which averages out to roughly 2% newly acquired per year. The authors' sentence is that adding a 3% annual cost decline to that puts half of physical work at cost parity in 2085. The rapid scenario raises both axes at once: quality-adjusted costs fall up to four times faster, and the rate at which robots gain new tasks doubles.

Assumption Baseline scenario Rapid scenario
Cost decline3% a year, all-in deployment cost12% a year hardware, 6% the rest
Capability growth2 percentage points of unexposed tasks become exposed each year4 percentage points a year
10% of tasks reach cost parity2060s (the main text says "about 40 years")around 2040
Half of physical tasks20852050
Janitors and cleaners20952050s

Taken from Appendix E and the main text. Even under the rapid scenario, the text notes, automating 90% of today's physical tasks takes 53 years.

Two Scenarios, Two Target Dates 2026 · today 2050 Fast scenario · 12%/yr 2085 Baseline scenario · 3%/yr
▲ Pebblous original diagram (scenario comparison reconstructed) — the widely reported 2050 is the fast scenario; the paper's baseline is 2085

The published side calculations reproduce exactly. A 3% annual decline reaching 20% in about 7 years and 70% in about 40 falls straight out of simple compounding. What does not reproduce is the headline dates, because the distribution of the gap between robot cost and labour cost across tasks was never published, nor was the order in which capability growth picks tasks off. This report uses the published values as given while recording that they cannot be independently reproduced.

The authors added a caveat of their own at this point. Because the scenarios apply one uniform decline rate and one uniform capability growth rate to every task, tasks reach cost parity in exactly today's cost order. Reality may not work that way, they say, and they offer an analogy for it: humanoid demonstrations show housework, while manufacturers may well pick the tasks useful in their own factories first, much as AI companies looked after coding agents first.

In the sentence immediately after that caveat, the authors tie both scenarios together in one line. How the rapid case should be read is written right there.

"Overall, robots would need to sustain record rates of price declines and quality improvements over the coming decades to enable rapid physical automation. Even so, job impacts are less certain with these cost projections, since they set aside forces like preferences and regulation as well as feedback effects like falling wages."

Two things are in that sentence. One is that the authors themselves called the price declines the rapid scenario requires "record" rates; before opening those grounds one by one below, it is fairer to put their own word down first. The other is an admission that the cost projections run with preferences, regulation and wage feedback set aside. A study that put the share of physical tasks blocked by regulation at 14% left that 14% out of the arithmetic when it produced its arrival dates.

5.2The One Pillar Under the 12% a Year

Every piece of material for the objections below was written by the authors in their own footnotes. Nothing is concealed. The issue is that the footnotes vanished from the headline.

2050 hangs on one assumption: hardware prices fall 12% a year. The historical grounding Appendix footnote 49 offers for it is a single quality-adjusted price index the International Federation of Robotics built for 1990 to 2005. Five caveats attach to that index.

# Caveat What was verified
1The index stoppedThe list-price index ended in 2005 and was never rebuilt. For 21 years there has been no index that could confirm or refute this assumption
2Three base years move half of it1990–2005 works out to 10.2% a year, but the body of the very paper Anthropic cites says that narrowing to 1993–2005 gives a decline of around 50% over 15 years. Annualised, that is 5.6%
320% of the quality adjustment is a computerThe federation's quality adjustment treats 20% of a robot's marginal cost as the controller, that is, a computer. The curve footnote 47 excluded by saying "robots are closer to cars than to computers" sits inside the rapid scenario's source index at a weight of 20%
4Small sample, list pricesTen models tracked per year, collected as list prices rather than transaction prices, across six countries
5Another index over the same period disagreesFootnote 49 says so itself: the Bank of Japan's quality-adjusted index fell only 3% a year over the same period

Items 1, 2 and 4 verified in the body and online appendix of the Graetz and Michaels paper Anthropic cites; item 3 in the federation's methodology notes; item 5 in Anthropic's footnote 49 itself. The annual rate in the second row is back-calculated from the paper's own wording.

Read the fifth row again. Two indices of the same kind of machine over the same period diverge at 10.2% a year and 3% a year. Anthropic took the larger one and disclosed the choice in a footnote. Disclosure is to its credit. That the entire rapid scenario stands on that one choice is the part to point at. Inside a single company, too, the two scenarios rest on different answers. The baseline's 3% comes from a judgment that robots are closer to cars than to computers, and the rapid scenario's 12% comes from an index that draws a fifth of its quality adjustment from computer prices. Not quite a contradiction. Scenarios are supposed to depict different worlds. But the two scenarios rest on different answers about what kind of good a robot is, and the authors wrote that answer down only on the baseline side.

5.33% a Year Already Assumes Robots Learn as Well as Solar

Looking at why prices fall through the lens of learning curves makes the distance between the two scenarios sharper. The rule of thumb is that unit cost falls a fixed proportion for every doubling of cumulative production, and solar modules are the standard case, having shown that proportion at around 20%. Six days before this study came out, the IFR's 2026 count put the global operational stock of industrial robots at 5,079,000, up 9% in a year, with annual installations at 603,307, up 11.2%.

Putting those two into the learning-curve equation gives an interesting result. A solar-grade learning rate of 20% applied to actual volume growth of 9% yields a decline of 2.7% a year, almost exactly the baseline scenario's 3%. Which is to say that 3% a year is not a conservative number; it already assumes the robot industry learns as well as solar did. So what would sustaining 12% a year require? At a 20% learning rate, cumulative volume would have to grow 48.7% every year, and even at the most generous 35% rate it would have to grow 22.8%. The actual figure is 9%, and the federation's own 2029 outlook is 7.2% a year. Hold 12% for 24 years and hardware ends up at one twenty-first of its price.

Precedent is not absent. Solar modules, semiconductors and batteries all did it. The shared conditions were mass production of modular identical units and compound growth in cumulative volume. New car prices, by contrast, were close to flat from 1990 to 2019 even after quality adjustment. And solar itself was never smooth: it stalled from 2005 to 2009 and actually rose from 2021 to 2023. "An average of 10% a year over fifty years" smooths over two plateaus and one reversal.

The most recent actual measurements point the other way. The Association for Advancing Automation counted 36,766 robots ordered in North America in 2025 for $2.25 billion, with units up 6.6% and dollars up 10.1%. Average price per unit rose 3.3%. Anthropic's footnote 47 says the same thing: prices have risen in recent years. That series is not a quality-adjusted index, though, so it cannot serve as a refutation of the 12%. All it supports is the smaller claim: no recent measurement backs 12% a year.

5.4The Forecasts That Reach 2050 Converge

Footnote 49 cites three investment-bank forecasts. Anthropic's description of them as the fast end of market forecasts is accurate. The problem lies elsewhere. Separated by the horizon each one reaches, the three take on a different shape.

Forecast Annual decline Horizon it reaches
Bank of America13.5% → 9.4%Ends at 2030, 2035 at the outside. Component-cost basis, assuming Chinese production
UBS5.8%2045–2050 (primary source not verified; via press coverage)
Morgan Stanley5.2%2050 (primary source not verified; via press coverage)
Anthropic rapid scenario12.0%2050

The Bank of America figures were verified in the March 2026 report itself. What Anthropic's footnote cites is the same institution's 2025 document, which is an image-only file we could not open, so the verified claim extends only as far as the March 2026 report reproducing the footnote's "about 13%" and "about 9%." The UBS and Morgan Stanley figures are recorded as forecasts only, their primary documents unverified.

The two forecasts that reach 2050 say somewhere between 5.2% and 5.8% a year, less than half of Anthropic's assumption. The only forecast that reaches 12% ends in 2035 and decelerates even within that span. A quiet convergence point also turns up here. The 5.6% from the previous section, obtained by moving the base year forward three years and remeasuring, meets the band of the two forecasts that run to 2050 at the same spot.

No conclusion that 2050 is impossible comes out of this. 2050 rests on an assumption that robots turn into a kind of good they have not once behaved like over the past 35 years, and the only historical grounding for that assumption is a single index that stopped 21 years ago. Of the six obstacles the Bank of America report lists in that same document, two run straight into this assumption. Today's humanoids are too fragile and too maintenance-heavy for round-the-clock industrial use, and repairs are expensive.

5.5Four Cost Assumptions out of the Factory World Too

There are four assumptions on the cost side: a cost of capital of 8%, a 10–12 year service life, hardware at one third of total cost, and one technician supervising five robots' worth of work. Something fair belongs first. These four were not presented as global constants; they came out of a single meat-processing robot case, only the hardware share carries an outside citation, and the appendix states explicitly that the supervision ratio is Claude's estimate. The arithmetic itself also checks out. Amortise an upfront $285,000 at an 8% cost of capital over 12 years and the annuity is $37,818, matching the main text's "about $38,000."

Checked against outside sources, the four do not stray far from consensus. An 8% cost of capital is close to the 7.7% weighted average cost of capital for listed machinery firms. The firms employing packers, meat processors and cleaners, though, are small and mid-sized businesses rather than listed corporations, and their cost of capital runs 10–15% with equipment loan rates of 8–18%. That direction biases robot costs low. Hardware at one third sits in the middle of the industry's consensus range.

Service life is where it diverges. Ten to twelve years matches the 12 the IFR uses to count operational stock, and for fixed robots inside a factory it may even be conservative. In the tier this study is new about it reverses. Commercial cleaning robots last around five years according to vendor material, and Morgan Stanley's humanoid model uses a six-year replacement cycle. Anthropic took the price path from that model and not the lifetime assumption, and that direction also biases robot costs low.

So this study's cost estimates are not sloppy. The assumptions come from the world of fixed industrial robots and they fit that world well. What this study is new about is warehouses and roads, and there the same assumptions bend the other way. It catches on exactly the same hook as section 4's "what was validated is the factory interior."

6

Why This Matters to Pebblous

Everything up to here gathers into one shape. A top-tier model was sent on 650,000 web searches to measure the future of physical AI, and the central uncertainty in the conclusion came not from the model's reasoning but from the nature of its inputs. That is why Pebblous spent so long with this report.

6.1Traces of the Claim That Data Is the Bottleneck

This study does not argue that data is the bottleneck in physical AI. It shows it, through the traces an attempt to measure that bottleneck left behind. With no record of how often a robot succeeded in the field and how often it called a human, vendor sentences about their own products filled the gap. The unstructured-environment tier staying at 2% of physical tasks can be read as a result from the same place. Where little data has been collected in an environment, there are few robots to show doing work in it.

6.2Evidence Tiers Recorded, Rating Stability Not

From a data quality standpoint, what this study has to teach is not a criticism but a practice: tiering the evidence. The authors recorded deployment, commercial claim and demonstration separately, gave transfer-based ratings their own field for the reason, kept the teleoperation flag, and published the full prompts and the dataset. That is why outsiders could recompute the sensitivities, and even identify the inequality sign they actually used. Writing the evidence tier into the schema is the best part of this work.

The limit sits in the same place. The quotations are truncated at five leading words, so what a reader can verify is not the source's sentence but the model's summary of it. The decision line runs between intermittent assistance and continuous operation, and the published fields never say which side of the line a rating sits on. The prompts promised a human reviewer audit three times, and neither an agreement rate nor the variance across repeated runs was reported. Evidence tiers were recorded; rating stability was not. Using model-produced estimates as an index requires both.

That lesson comes with a price tag attached, and the paper Anthropic cites in its own appendix footnote already supplies the figure. To use model-made labels in estimation, relabelling 5% of them by hand is enough to restore the confidence intervals. Five per cent of 7,594 physical tasks is about 380. A peer-reviewed study in the same genre went as far as recommending in writing that researchers use at least two models and report inter-model agreement. The concrete conditions for data fit to train on are right here. A model-made label has to arrive with a small human sample attached, one that measures that label's error. Without it, no amount of labels will hold a downstream inference up.

6.3Three Questions for Next Year's Vendor Deck

Customers in manufacturing, logistics and construction will receive material next year citing this study's headline. The moment "robots can do 74% of our tasks" lands in an internal report, a practitioner has three questions to ask. Which environment does that 74% presuppose, and is there a plan to actually rebuild our workplace? Is the evidence behind that rating a robot installed on our floor, or a robot doing some similar task somewhere else? Does "can do" mean "repeats at a quality our line will accept"? All three are questions this study already answered in its own appendices, and all three have the same shape as the questions used to write data acceptance criteria.

6.4The Empty Seat

The configuration fits in one line. The people building robots write down what their products can do. The people building models read that writing and estimate an economy from it. Both sides worked in good faith, and between them the seat for a number counted in the field is empty. Data gathered in environments nobody controlled, and a standard for judging whether that data is good enough. That is where Pebblous stands. We wrote this report the way we did because letting a reader see the empty seat for themselves beats selling them a conclusion.

One job here is still undone. All five studies that have measured grader reproducibility so far concern language-model exposure indices, and no study running robot environment tier ratings through two models could be found. With 18,796 dataset rows and the full prompts public, rescoring part of the 7,594 physical tasks on a different model and producing an agreement rate is genuinely feasible. That is an observation that the job is available to anyone, not an announcement that Pebblous will do it. Why data gathered in uncontrolled field conditions is a different class of asset is something we set out once before, in a piece on rollator navigation data.

Sections 1 through 5 report what was verified in the published main text, the appendices and the public dataset; section 6 is the part those documents do not cover. Please read them separately. The verbatim quotations in the body were checked directly against the original main-text and appendix PDFs, and every figure attributed to a Pebblous recount comes from aggregating the public dataset ourselves. The headline 74% and 34% are employment-weighted and could not be reproduced from the public files alone, which is noted in the body where they appear. As of 3 October 2026, three days after release, no academic rebuttal to this report has been identified. Thank you for reading this far.

R

References

Sources are grouped because they differ in standing. The first group is the primary material of the report this article examines, and every verbatim quotation and figure in the body was checked directly against it. The second is the public data that report drew in from outside. The third is academic literature whose original text or tables we opened ourselves. In the fourth group, industry material is marked line by line according to whether primary verification succeeded or the figure came via press coverage. Every value in the body introduced with "Pebblous recounted" comes from aggregating datasets 1–4 directly.

The report under examination (checked against the primary text)

  • 1.Russell Legate-Yang, Maxim Massenkoff. "What work can robots do?" Anthropic, 30 September 2026. anthropic.com — a company research report rather than a journal article. There is no arXiv number, no DOI and no peer review, and the citation format the authors supply is the one used for online documents. Hence "report" rather than "paper" throughout this article.
  • 2.The same study's main-text PDF (45 pages). The four headline numbers, the tier shares, the backtest results, the arrival dates for both scenarios, and footnotes 43, 46, 47 and 49 come from here.
  • 3.The same study's appendix PDF (57 pages, A–F). Most of this article's argument lives here. The three performance conditions and the cap and counterfactual rules are in Appendix F, the weighted majority formula in A.3, the transfer-exclusion sensitivity in A.4, the 42 percentage points in C.2, the backtest caveat in B.3, and the cost assumptions in D.1.
  • 4.Anthropic. EconomicIndex / robot_exposure dataset, CC BY 4.0. Hugging Face — 18,796 task rows and 923 occupation rows. The 7,594 physical tasks, the 56,933 citations, the 105 E3 tasks, the inequality-sign reproduction, the transfer-exclusion recount (66.5% → 49.3%), the domain tally (30.8%) and the teleoperated-or-unclear share of determining evidence (23.6%) all come out of this file.

Public data the report used

  • 5.O*NET 29.3 Database, US Department of Labor Employment and Training Administration / BLS Occupational Employment and Wage Statistics (OEWS 2025) / BLS Employer Costs for Employee Compensation (ECEC 2026) / US Census Bureau American Community Survey 2020–2024 / BLS Occupational Requirements Survey (ORS 2023, 2025) — the last of these is the comparison data for the classification validation in Appendix A.2.

Academic literature cross-checked

  • 6.Mingyu Yin, Hieu Vu, Claudia Persico. "How (un)Stable Are LLM Occupational Exposure Scores? Evidence from Multi-Model Replication." NBER Working Paper 35110, April 2026. DOI 10.3386/w35110 — section 3's cross-model agreement of 56.9–73%, the 3.6× spread in occupational means and the sign flip in downstream regressions were verified in the abstract and body; the within-model repeat agreement of 90.8–99.0% in appendix Table A2. The model used there is Claude 4.5; the report uses Opus 5.
  • 7.Jens Ludwig, Sendhil Mullainathan, Ashesh Rambachan. NBER Working Paper 33344, January 2025 (revised December 2025). DOI 10.3386/w33344 — the very paper Anthropic cites in appendix footnote 10. The distinction between the conditions for prediction problems and estimation problems, and the simulation result that relabelling 5% of the labels by hand restores confidence interval coverage, come from Table 3.
  • 8.Tyna Eloundou, Sam Manning, Pamela Mishkin, Daniel Rock. "GPTs are GPTs." arXiv:2303.10130 / Science 384(6702), 1306–1308 (2024). DOI 10.1126/science.adj0998 — human-model agreement of 65.6–82.1% and the 24% label flip on a rubric wording change were verified in Table 2, and the statement that human annotations are the primary results in section 3.4.2. ⚠️ Version note: the working paper abstract gives headline figures on the model labels, while the published abstract gives them on the human labels.
  • 9.Emilio Colombo, Fabio Mercorio, Mario Mezzanzanica, Antonio Serino. IJCAI 2025, art. 1066 / arXiv:2407.19204 v3 — a peer-reviewed case validated with three open-weight models and three human annotators. The verbatim statement giving reproducibility as the reason for that model choice comes from here.
  • 10.Schaal. arXiv:2510.13369 (2025) — cross-model correlation of 0.60–0.81. ⚠️ This study reads figures at that level as robust while reference 6 reads them as fragile. The accurate characterisation is not that academics distrust model graders but that they publish numbers and argue about them.
  • 11.Carl Frey, Michael Osborne. Technological Forecasting and Social Change 114, 254–280 (2017). DOI 10.1016/j.techfore.2016.08.019 / Michael Webb. "The Impact of Artificial Intelligence on the Labor Market." Stanford Working Paper (2020) — the first two rows of the lineage table in section 3. Journal publication of the latter could not be confirmed.
  • 12.Daron Acemoglu, Pascual Restrepo. "Robots and Jobs." Journal of Political Economy 128(6), 2188–2244 (2020). DOI 10.1086/705716 — section 4's 0.2 percentage points and 0.42% are the published values and differ from those in the earlier working paper. The exclusion rule removing automated storage and retrieval systems from the count was verified in a working paper footnote.
  • 13.NIST SP 1011-II-1.0 (ALFUS, 2007) · ISO 8373:2021 · SAE J3016 APR2021 · RLBench and other manipulation benchmarks — the counterexamples in section 3.6. Three of these are taxonomies of design intent rather than of performance, and the one precedent carrying a number is the 2007 document's "approximately 50% dependence on operator input."
  • 14.Rivière, Denain. Epoch AI (February 2026) — the analysis finding robots three to ten times slower than people. Anthropic cites it three times, and footnote 43 of the main text says a cleaning robot is up to 10× slower than a human.

Industry and statistical sources

  • 15.Georg Graetz, Guy Michaels. "Robots at Work." Review of Economics and Statistics (2018). Body and online appendix — both verbatim quotations in section 5 come from here: the sentence that the list-price index stops in 2005, and the sentence that narrowing to 1993–2005 still gives a decline of around 50%.
  • 16.International Federation of Robotics. World Robotics (2006), Annex C methodology — the basis for treating 20% of a robot's marginal cost as the controller in the quality adjustment. Verified as reproduced in the appendix to reference 15.
  • 17.International Federation of Robotics. World Robotics 2026, released 24 September 2026. ifr.org — operational stock of 5,079,000 (+9%), annual installations of 603,307 (+11.2%), 7,000 full-size humanoids across 174 manufacturers, and the verbatim statement that humanoids remain between the lab and the pilot.
  • 18.BofA Institute. Transformation — Physical AI, part 2: Humanoid robots, 12 March 2026 — the deceleration from 13.5% to 9.4% a year on a component-cost basis, and the six obstacles. ⚠️ What Anthropic's footnote 49 cites is the same institution's 2025 document, an image-only file we could not open. The verified claim extends only as far as this later document reproducing the footnote's "about 13%" and "about 9%."
  • 19.UBS (2025) and Morgan Stanley (2025) humanoid price forecasts — ⚠️ primary documents could not be verified, so these came via press coverage and summaries. The 5.2–5.8% a year is recorded as a forecast rather than asserted. Morgan Stanley's six-year replacement cycle carries the same standing.
  • 20.A3 (Association for Advancing Automation) North American robot orders for 2025, 36,766 units and $2.25 billion / Agility Robotics SEC Form S-4/A (FY2025 audited total revenue of $1.78 million) / Tesla SEC 10-Q (Q2 2026 installed capacity table) / Damodaran industry cost of capital data — ⚠️ the A3 figures are not a quality-adjusted index and were not used to refute the 12% a year.

Secondary coverage

  • 21.Michael B. Kelley. Yahoo Finance, 1 October 2026. finance.yahoo.com — the article section 1 uses to separate what the coverage carried from what it left out. Missing from it are 2085, the E0–E3 tiers, the appendix transfer-exclusion sensitivity, the 42 percentage points, and the 14% blocked by regulation.

Related Pebblous articles