Executive Summary

AI agents hire people to do things in the physical world and pay them automatically once the job is confirmed. Six months after a marketplace called RentAHuman opened in February 2026, someone has counted what those public listings actually ask of a worker. The audit comes from Iman YeckehZaare of the MIT Center for Collective Intelligence, posted to arXiv on August 19 and accepted to the ACM HCOMP 2026 conference in September.

A one-day snapshot on May 31 returned 981 records, of which 779 advertise a task for a person to complete and form the analysis population. Two coders labeled thirteen features by hand and scored each listing from 0 to 5. On that scale 438 listings landed in the top two bands, at 4 or above. What those 438 actually demand, though, splits into 154 distinct combinations, and the ten most common of them cover only 155 listings between them. The author's conclusion is that the checklist, not the score, tells a worker what a listing entails.

What this audit measures is what listings advertise, not what workers submitted, were paid, or experienced. It is a single point in time from what is effectively a single platform, and the author does not generalize past it. This article reads the findings inside that fence, separating what the audit confirmed from what it did not.

Key Figures

Source: YeckehZaare, arXiv:2608.18547 (2026-08-19), based on the 779 eligible listings

56.2%

Listings scoring 4 or 5

438 of 779. The author's name for this band is severe, the heaviest on his scale

154

Requirement combinations in that band

The ten most common cover 155 listings (35.4%), so no profile dominates

45.4%

Identity-linked group among score 4 or 5

199 listings. The largest of nine non-overlapping groups, twice the size of the next

10.4%

Identity-proof listings caught by automated rules

Out of the 241 that human review finally confirmed. This judgment does not automate yet

1

The Price Is Posted. The Conditions Are Not.

The RentAHuman homepage is blunt. "AI needs your body," it reads, and adds that "AI can't touch grass. You can." The site launched on February 1, 2026 with 130 people signed up on day one, a thousand the next day, and another 145,000 a day after that. Those growth figures are the platform's own. Futurism, looking at the site two days after launch, saw a claim of more than 73,000 registered people while only 83 profiles were actually visible on the browse tab. In the platform's vocabulary the people are meatworkers, the jobs agents post are bounties, and the photo or geolocation that triggers automatic payment is proof of presence.

A developer connects an agent to the platform's MCP server or REST API, and the software scans the human directory by GPS coordinates and hourly rate, then messages someone directly or posts a public bounty. Payment sits in escrow until the agent confirms completion, at which point it releases in crypto or through Stripe. By the numbers Built In gathered in March, the site had logged more than four million visits and was approaching 600,000 registered meatworkers, against roughly 11,300 active bounties and about 5,500 jobs completed to date. That is close to fifty workers for every available task. A Wired reporter dropped his rate to $5 an hour and still heard nothing back, and other coverage has pointed to a flood of low-value tasks and listings that look like crypto scams.

How RentAHuman Hands Off Work AI agent Needs a task MCP / API link Search by GPS, rate Post a bounty DM or public listing Prove completion Proof burden here Auto payout Crypto or Stripe What this paper counts is step four — what a person must hand over to prove the job is done
▲ Original Pebblous diagram, reconstructed from the paper's account of the transaction flow

This paper does not relitigate the spectacle. It takes a position that is neither defense nor exposé and counts what the public listings ask of a worker. The concept it introduces is proof burden. The premise is that completing a task and making the completion verifiable are two different jobs.

One sentence in the introduction carries the whole study. "A posted price advertises an amount or rate; it does not say everything a worker may have to reveal, record, do in public, or provide again later." When a listing asks someone to reveal identity or location, post from a personal account, or travel somewhere solely to photograph the result of an otherwise remote task, then, in the paper's words, "proving the work can itself become a source of exposure."

2

Thirteen Boxes, Counted by Hand

Comparing free-text listings means first turning them into something countable. The author settled on thirteen features: eleven kinds of evidence (text, link, screenshot, photo, video, identity, account, location, phone, financial, public post) plus recurring monitoring and physical-world action. Recurring monitoring covers work, availability, or evidence demanded again at a later point. Physical-world action marks cases where the act itself creates exposure even when that act is not submitted as evidence.

Reliability comes from the procedure. Two newly recruited coders independently labeled all 981 records, and a blinded adjudicator, who could see both coders' answers but no superseded labels, resolved 873 of them. Listings where the two coders had already agreed on everything were routed to the adjudicator as well, which was deliberate. On the 0 to 5 score the two coders matched exactly on 81.8% of records, and on the binary question of whether a listing reached 4 they agreed on 90.2%.

Adjudication did not overturn the headline. Before it, the two coders had classified 59.2% and 58.5% of eligible listings as severe, against a final 56.2%. What leans more heavily on the adjudicator is not the overall share but which feature put a listing in its band. The coders disagreed most often on text proof (82 listings), recurring monitoring (62), and identity (61), and the latter two are precisely the features this audit's conclusions rest on. On contested labels the adjudicator sided with each coder almost equally, 55% against 45%, so the final labels do not simply mirror one person's reading.

Two populations need keeping apart. Searches returned 981 records in total (980 from RentAHuman, one from a similar market called Human Pages). A preregistered content screen removed promotional posts, service offers, and records requesting no human action, leaving 779 listings as the primary population. Every prevalence figure and score share below uses 779 as the denominator. Already-cancelled listings were kept and counted as advertised requests.

There were a lot of those. Of the 779 eligible listings, 404 were already cancelled at collection (531 of the 981 returned). The author's reasoning is that cancellation does not change what a listing asked of a worker. Collection has gaps too. Some public searches failed, and every query filtered to the recruiting category failed. A broad search's first page showed one recruiting listing, so the author notes that more may be missing from that category.

One coding rule shapes the character of the results. A word like "verify" or "photo" appearing in a listing was not enough to mark a feature. It had to be tied to completion. Location proof means a demand to demonstrate where you actually were, which the codebook separates from the mere fact of having to travel somewhere.

Here is how often each of the thirteen features turned up. The three marked in orange are the ones that return in the next section.

What 779 listings asked for Share (listings) Text proof 57.8% (450) Physical-world action 48.5% (378) Photo proof 37.1% (289) Identity proof 30.9% (241) Link proof 29.7% (231) Video proof 26.6% (207) Account proof 24.4% (190) Screenshot proof 20.3% (158) Recurring monitoring 19.4% (151) Public-post proof 15.8% (123) Location proof 10.7% (83) Financial proof 10.3% (80) Phone proof 5.5% (43) Orange marks the three features that drive the score-4-or-5 band and the comparison in the next section
▲ Figures from arXiv:2608.18547 Figure 2, redrawn by Pebblous. Denominator is the 779 eligible listings

Location proof landing at 10.7% has a backstory. The first coding round put it at 25.9%, having conflated location with physical-world action. Once the second round separated the definitions, the figure fell by more than half. Traveling somewhere and proving you were there turn out to be different demands.

3

Six in Ten Landed in the Top Two Bands

The Proof Burden Score is assigned in two steps. A coder first picks a base tier from 0 to 4. No specified proof or a minimal text acknowledgement is 0. Private text about the work, or a private link to such text, is 1. A screenshot or photo of the work product is 2. Video, public-post, account, location, or financial proof is 3. Tier 4 is identity proof, or recurring monitoring, or physical-world action paired with phone proof, with public posting or account use, or with location proof plus photo, video, or financial evidence. Three modifier conditions are then checked, each adding a point, and the total is capped at 5.

The counts behind those modifiers explain why scores pile up at the ceiling. More than one evidence type was requested in 551 listings, physical-world action was combined with proof or monitoring in 360, and public posting or account use was required in 200. A listing already at base tier 4 needs only one of these to hit 5.

The distribution leans high. 492 listings (63.2%) score at least 3 and 438 (56.2%) score at least 4. Score 5 alone accounts for 372 listings (47.8%), and that 5 is a ceiling rather than a measurement. For 222 of those 372 (59.7%), the recorded base tier plus modifier points exceeds 5. The heaviest band contains spread that the scale cannot show.

On precision the author reports two intervals. Treating every listing as unrelated to every other gives a 95% Wilson interval of 52.7 to 59.7% around the 56.2%. Grouping listings that share a requester display name into 469 clusters widens it to 51.5 to 60.9%. Both are labeled reference intervals, and neither generalizes beyond this nonrandom snapshot.

Shaking the scoring rules barely moves the figure. Applying the preregistered duplicate rule removes 42 eligible listings and shifts the severe share to 56.3%. Dropping the modifier point for public posting or account use altogether pushes only 15 of the 779 listings below 4. Six in ten is not a number balanced on one or two rows of the rubric.

What the paper calls more important is the number that follows. The 438 listings scoring 4 or 5 contain 154 different combinations of the thirteen features, and the ten most common cover only 155 listings, 35.4% of the band. No single profile dominates. Two listings can both score 4 while one wants your identity and the other wants a fresh photo every day.

To summarize that variety, the author assigns each score-4-or-5 listing to the first qualifying group in a fixed order, once only. The result is below, and the groups do not overlap.

First qualifying group Listings Share of score 4 or 5
Identity-linked proof 199 45.4%
Physical-world action plus proof 102 23.3%
Recurring monitoring 101 23.1%
Account proof 17 3.9%
Remaining five groups (video, financial, public post plus account use, location plus financial, other) 19 4.3%

▲ arXiv:2608.18547 Table 3. Nine non-overlapping groups, condensed into five rows

The largest row in the table is identity-linked proof. Gig work is built on the premise of working briefly and anonymously and moving on, and here the evidence of having done the job converges on who you are. By category, score-4-or-5 shares run highest in delivery errands, events and social, marketing campaigns, and hiring, while documentation and writing content hold the largest shares of score 0.

4

Did Agent-Posted Listings Differ?

This is the question most readers will want answered. It is also the part of the paper that needs the most care. Platform metadata labels some requester accounts as an agent or a bot. Seventy-two listings carry that label and 676 carry a human label.

On the score itself the answer is flat. The share scoring 4 or 5 is 58.3% among agent-or-bot-labeled listings against 55.9% among human-labeled ones, an odds ratio of 1.10 with a 95% reference interval of 0.56 to 2.19 (p = .778). In the author's words, "we therefore find no clear group difference in the share scoring 4 or 5." Removing the agent-heaviest mixed cluster actually reverses the odds ratio to 0.70, and removing every mixed name drops it to 0.44. The 72 agent-labeled listings come from only twenty display names, which means twenty independent units of observation.

Individual features show a sharper pattern. Agent-or-bot-labeled listings have higher observed shares of physical-world action (75.0% against 44.7%), recurring monitoring (38.9% against 17.3%), photo proof (56.9% against 35.1%), and financial proof (13.9% against 9.5%). Location proof runs the other way (2.8% against 11.7%), and link and video proof are also lower in the agent group. Not every difference points the same direction.

Individual Features Split by Agent vs. Human Labels Agent-labeled Human-labeled Physical-world action 75.0% 44.7% Recurring monitoring 38.9% 17.3% Photo proof 56.9% 35.1% Financial proof 13.9% 9.5% Location proof 2.8% 11.7% Only location proof runs higher for human-labeled listings. The author reports none of these individual gaps as confirmed
▲ Original Pebblous diagram, reconstructed from arXiv:2608.18547 body figures. Individual-feature comparison; test calibration is unstable

None of those individual gaps is confirmed either. The author built simulated datasets in which the two groups genuinely did not differ, then checked how often his test claimed a difference anyway. For the three-feature test the false positive rate came in at 3.7 to 7.7%, near the intended 5%, but rarer features fared badly, and phone-proof tests reached 41.5% false positives. Results also shift with the sample. Removing the single name that supplies the most agent listings reverses the recurring-monitoring gap to 3 of 36 (8.3%) in the agent group against 110 of 663 (16.6%) among human-labeled listings. Hence the paper's line: "because the individual-feature tests can be poorly calibrated, none confirms a single-feature difference."

The order of what the author did next matters. After seeing these differences, he defined an exploratory three-feature measure marking any listing that requests at least one of physical-world action, location proof, or recurring monitoring. That measure appears in 54 of 72 agent-or-bot-labeled listings (75.0%) and 374 of 676 human-labeled ones (55.3%). It is an eminently quotable pair of numbers, and the fact that the comparison was chosen after the data had been examined has to travel with it. Cutting the same measure finer, at least two of the three features appear in 41.7% against 17.2%, and all three in 0.0% against 1.2%.

The author fences the result four ways in his conclusion: the comparison was created after seeing the data; the agent-or-bot-labeled listings were posted under only twenty display names; category mixes differ between the groups; and the requester labels are self-reported or platform-assigned, so they cannot establish who designed or supervised a task. His verdict reads: "This is a hypothesis, not a confirmed difference." Or, in the summary line of the results section, "no difference is confirmed, and the consistent three-feature pattern remains post hoc."

The category problem in particular is concrete. Research fieldwork accounts for 41.7% of agent-or-bot-labeled listings against 7.8% of human-labeled ones, and all 30 of those agent-labeled fieldwork listings meet the three-feature measure. A Mantel-Haenszel odds ratio pooling the within-category comparisons is 3.9 (reference interval 1.9 to 8.2), and omitting research fieldwork reduces it to 1.9 (0.8 to 4.3). Whether agents ask for more, or simply post more fieldwork, is not something this data separates.

5

What Machines Cannot Read, People Prove

The paper also reports an attempt to automate its own method. The author wrote a deterministic extractor of keyword and regular-expression rules, with no machine learning and no generative AI, and used it only to surface candidates. It never touched the final labels, and neither the coders nor the adjudicator saw its matches. Checked afterwards against the human labels, it performed poorly where it mattered: the rules correctly flagged only 10.4% of the listings finally labeled identity proof and 13.9% of those finally labeled recurring monitoring.

Missing things is only half of it. The score implied by the automated matches differed from the final score for 373 of the 779 listings, and the errors ran in both directions. Identity proof went from 44 rule matches to 241 final labels and recurring monitoring from 33 to 151, while photo proof fell from 387 to 289, public-post proof from 210 to 123, and location proof from 169 to 83. A word appearing does not mean the requirement is there, and its absence does not mean it is not. Close to half the corpus still needs a person to read it.

That reading has a price. Two coders each read all 981 listings and the adjudicator resolved 873 of them, and all three were paid $37 per hour on average. Pinning down what a single listing demands took two to three passes of human attention. The cost of verification does not disappear; it attaches to someone. In this audit a research budget paid it. In the market, the worker does.

Two numbers overlap here. The two most exposing demands happen to be the two hardest to detect automatically. Identity proof makes up 45.4% of the score-4-or-5 band while being visible to the machine in about one case in ten. The further automation advances, the more both jobs stay on the human side: reading what was demanded, and satisfying it.

Anyone who has designed a data labeling or verification pipeline will recognize the shape of this. What has to be submitted before a system will believe the work is done is a decision made daily. Requiring a screen recording where a screenshot would do, attaching identity checks to anonymous workers, re-verifying an already-verified result on a schedule: each of these buys trust. The price is paid out of worker exposure rather than the requester's budget. What makes this paper useful is that it split that payment into thirteen boxes and counted them.

So did people avoid the heavy listings? The author asked the same question. Application counts are visible on 404 of the 779 listings (51.9%), and higher scores were weakly associated with fewer applications per position. Adding how long a listing had been public attenuates that association to null, and because time on the market is confounded with the visible application count, the author declines to interpret it. A ten-day follow-up found no clear difference in continued public visibility between severe listings and those below 4 either. The intuition that heavier demands drive workers away is not something this data confirms.

The author also draws a line around how far the tool may be used. The thirteen-requirement checklist and the score "are only for manual research on advertised requirements; they should not yet guide workers, task ranking, pay, moderation, or enforcement." There is no worker-validated measure and no usable automated detector yet. If classifiers are used at all, the recommendation is that they be used "only if workers endorse the categories, with human review whenever a classifier is uncertain."

Editor's Note: The scene Pebblous keeps meeting in data quality work has the same shape as this audit. Requests to compress quality into one composite score arrive constantly, and the score does work well on a dashboard. The trouble starts when two datasets with the same score have their defects in entirely different places. That is what the number 154 is saying. Use the score to sort, and read the checklist to decide what to fix.

A second study looked at the same platform from another angle in February. Pulak Mehta analyzed 303 RentAHuman bounties and found that 99 of them (32.7%) originated from programmatic channels such as API keys or MCP, and identified six active abuse classes, including credential fraud and identity impersonation, purchasable for a median of $25 per worker. One audit counts what the market asks of people; the other counts what the market is being used for. Both point the same way. The market where agents hire humans has stopped being a talking point and become something measurable. The paper is at arXiv:2608.18547.

R

References

Academic Papers

Press Coverage