Executive Summary
Figure AI unveiled Index on August 25. It is an app that buys, straight from the people who shoot it, the video used to train Helix, the AI stack that runs the company's humanoid robots. The reason sits in the first line of the announcement. The data needed to scale a truly general purpose robot does not exist on the internet, and it has to come from the real world. Figure says it tried buying data first, that vendors could not hit the bar Helix requires, and that it therefore built its own route to collection.
At launch the counters read 264,000 app downloads, 16 million videos, 108 countries, and $15 million paid out to Creators. The upload rate rose from 30 minutes of video every second in late August to 35 minutes per second in the first week of September, which works out to 5.7 years of human work arriving each day. What a Creator earns per minute, and how many of those videos survive five stages of review to reach training, appear nowhere in the company's materials.
Sections 1, 2, and 4 follow only what Figure's announcement, its Index page, and its privacy policy say. Section 3 and the second half of Section 5 are this article's own reading of the launch as a data procurement problem.
Key figures
Source: Figure, Introducing Index (August 25, 2026) and the Index page; Humanoids Daily (September 8, 2026)
16 million
Videos uploaded
Gathered across 108 countries. Figure published this on August 25, and the Index page still carried the same number on September 13
35 min/sec
Upload rate
For the week ending September 4, which is 5.7 years of human work a day. At the late-August launch it was 30 minutes per second, or 4.9 years a day
$15 million
Paid to Creators to date
Figure has committed over $1 billion to data and compute across the next 12 months. It has not said what one minute is worth
0
Published pass rates
Three of the five pipeline stages discard video, yet no source gives the share that gets filtered out
Figure Buys the Scenes the Internet Does Not Have
Language models learned by scraping text that people had already written. A robot has to learn something else: folding laundry, moving dishes, twisting a handle. That footage is not piled up on the web. Figure's explanation starts from the same place. Scaling a general purpose robot takes a global sampling of physics captured across every environment on earth, and the web holds no such sample.
The company makes the same diagnosis about its own position. In the September 3 post announcing a compute supply deal, Figure wrote that it is "entering a phase where we are largely bound by data and compute needed to train Helix." That agreement covers up to 100,000 GPUs on the NVIDIA Vera Rubin platform, with initial deployment targeted for the second half of 2027 in Barstow, Texas, an initial commitment of $3.5 billion and stated intent to scale to over $6 billion. Getting hold of the material to teach the robot now sits ahead of building a better robot.
Buying the data came first. By the company's account, vendors could not hit the throughput, diversity, or quality bar Helix requires, so Figure spent four months building an app in stealth and released it on August 25 under the name Index. The structure opens on two sides. A Creator who applies and is approved receives a recording device, films a day of ordinary work at home or on the job, and is paid by the minute. On the other side, a household or a business books a Creator to come handle chores and errands, and the work leaves behind a first-person recording. Applying does not mean filming right away. The app listing says demand is currently very high, that an application puts you on the list, and that Figure will notify you as soon as it is your turn.
The tasks run from cooking, cleaning, and laundry at home to work inside logistics centers, restaurants, factories, and offices. The announcement says tasks as obscure as cleaning kitty litter, changing oil, and busing restaurant tables have come in, and that the company welcomes the diversity. One sentence in the closing paragraph sums up the direction of the business. Index is laying the groundwork for ordering robots as a service, and today a person comes to clean your house while eventually a robot will do everything for you.
Thirty-Five Minutes Arrive Every Second
The scale at launch was 264,000 app downloads, over 44,000 weekly active Creators, 16 million videos, and 108 countries. Uploads were being processed at 30 minutes of video every second, which the company put at 4.9 years of human work a day. The figures two weeks later are steeper. Weekly active users for the week ending September 4 came to 69,943, and the upload rate reached 35 minutes per second, or 5.7 years a day. On May 1 the same count stood at 99.
Figure pushes diversity as hard as speed. Per 1,000 hours collected, it counts 373 unique tasks, 1,146 unique manipulated objects, and 116 unique environments. Every new Creator brings an unseen environment, unfamiliar objects, and an idiosyncratic way of completing a task, the kind of long-tail variation the company calls nearly impossible to define upfront. Figure buys that variety by adding people.
How much this video actually captures has not been established. Humanoids Daily closed the same report by calling the data monocular smartphone video and noting that open questions remain about whether it can provide the fine-grained force dynamics and tactile feedback required for intricate physical assembly. Figure's own app listing says approved Creators are sent a recording device, so the two accounts split at the first question of what the footage is shot on. That report ends by saying that finding willing human contributors is not the bottleneck in Figure's equation.
The Rate per Minute and the Pass Rate Are Blank
From here this article reads the announcement again, through the lens of data procurement. Two blanks remain once the published numbers are set against the ones that were left out.
The first is the price. The Index page, the app listing, and the announcement all say only that Creators get paid by the minute. How much one minute pays is nowhere in them. Divided across 16 million videos, the $15 million paid out lands under a dollar per video, and that quotient is not a rate. Payment is keyed to recording time rather than video count, the videos run to different lengths, and the 16 million in the denominator includes footage that was dropped in review. From a Creator's side, there is no way to work out before starting what an hour of filming brings in.
The App Store listing tells applicants that while their application is under review they can "browse the task list, preview free services, and see how Creator earnings work." The number sits on a screen that opens only to people who have applied and been approved. Anyone who wants to weigh this work against other gig platforms stops in front of that screen.
The second blank is the share. Figure described its pipeline in some detail, across five stages. Automated filters screen first for technical, visual, and semantic quality. Human analysts then audit samples at the user level for deliberate attempts to evade those filters. To hold diversity, the company embeds each video segment and discards anything above a similarity threshold with previously accepted data. The remainder is rebalanced with task quotas and embedding-based clusters, and hierarchical text captions go on at the end, one set per surviving episode. Three of the five stages throw video away, and yet not one line says what percentage is discarded and what percentage reaches training.
Review results travel back to the Creator. The announcement says the data infrastructure was rebuilt around constraints more typical of a consumer app, listing 24/7 availability, continuous large-scale compute for processing, and "real-time feedback to users that helps improve future collections." A signal about what to fix reaches the person who shot the video before the next session. Figure has not released the aggregate, meaning how much of what came in was kept.
The results of all that collecting land in the same place. Figure has published a result before showing that human video carries over into robot behavior. Announcing Project Go-Big last September, the company said it had used "100% egocentric human video data, collected passively as people do behaviors in real Brookfield homes" to train Helix to translate human navigation strategies into robot control, with no robot demonstrations whatsoever. Figure has not yet reported how far the data Index gathered has pushed the model. The announcement says only that "the generalization results we're seeing internally are already validating this thesis, and we will be sharing more in detail on this soon." The number for how much was collected is out. The number for what the collection changed waits on a later announcement.
Speed is a metric that boasts about the numerator. The amount of data that actually reaches training is speed times pass rate, and the second factor has not been published. Nobody outside can work out how many of those 35 minutes a second land on the model, and the only party that knows the difference is the company receiving the data.
Only the Person Filming Signed Anything
The video coming into Index is somebody's living room, kitchen, and laundry room. The privacy policy Figure publishes, though, is one company-wide document, last updated January 21, 2026. That is seven months before Index opened its doors. The site's terms and conditions still carry a February 1, 2023 date.
One of the categories that document lists under collection is Sensory Data. The original text defines it as "photos, videos, or recordings of you and/or of your environment." The person filming is not the only subject; the space that person occupies falls inside the collection scope from the start. The section on disclosure to third parties carries this sentence: "Depending on state laws that may be applicable to you, some of these disclosures may constitute a 'sale' of your Personal Data." Later the same document says, "we do not sell, share, or process your Personal Data for the purposes of targeted advertising, and have not done so over the last 12 months."
The two sentences have to be read together. Figure draws a line at selling for advertising purposes while leaving open that disclosures to service providers and partners may be classified as a sale under state law. A stronger telling circulates in secondary summaries, that the policy lets the company sell video of the inside of a home to commercial buyers. The primary text supports only what is quoted above.
Read the collection table row by row and Sensory Data occupies a narrower place, not a wider one. Of the six categories the policy sets out, Sensory Data is the only one whose stated purposes leave out Marketing the Services, and the only one whose list of recipients stops at Service Providers. Profile and contact data, web analytics and device or IP data, and geolocation data each carry Advertising Partners alongside. The same document's Business Transfers clause, on the other hand, says all the Personal Data it collects may be transferred to a third party in a merger, acquisition, bankruptcy, or other transaction in which that party assumes control of the business. The retention section gives examples: first name, last name, email, the contents of a message sent through the contact form, and device or IP records. Video is not among them. On the question of how long the footage is kept, the heaviest item in the table is the one left outside the examples.
The real gap lies elsewhere. The person filming filled out an application and received equipment, while the family living in that home and the visitor who came by that day signed nothing. The privacy policy has no clause on the people who appear in the video alongside the Creator. The app carries a 17+ age rating, and that is the age of the person using the app, not the age of the people on screen.
The children's clause makes the mismatch plainer. The policy says "We do not knowingly collect or solicit Personal Data from children under 18 years of age" and asks anyone under 18 not to register for the Services or send any Personal Data. The child that sentence has in mind is a child sitting in front of a sign-up screen. In video of household chores the child is not the one sending information but the one caught at the back of the kitchen. A request not to register does not reach that position. Collecting human-perspective video inside homes did not begin with Index either. Announcing Project Go-Big last September, Figure said it was gathering video of human behavior with the cooperation of Brookfield's more than 100,000 residential units, and described that video as collected passively while people did their ordinary behaviors.
Korea Collects on Factory Floors, Not in Homes
Korea works the same bottleneck on a different stage. The Ministry of Science and ICT decided to build a data training center for physical AI training data and moved to secure a budget for it. The plan the Korea Economic Daily described on June 25 has two tiers. A central hub in Seoul and elsewhere produces and processes the basic behavior data and synthetic data a robot needs to operate, while regional AI transformation hubs in provinces such as North Jeolla and South Gyeongsang gather data from manufacturing, agriculture, logistics, and service sites. The unit of collection is concrete as well. Motions that break down into reaching, grasping, twisting, and pulling, as in picking fruit, get gathered first and serve as the base layer for robot learning. Underneath sits the physical-AI-based AI transformation R&D program in South Gyeongsang and North Jeolla, which runs to 676.3 billion won for South Gyeongsang and 736.8 billion won for North Jeolla, 1.4131 trillion won in total through 2030.
That article also states the rule for what comes first. The plan is to draw on the sites of small and mid-sized manufacturers, regional demonstration projects, and data startups already hold, gathering lower-sensitivity basic behavior data first and filling the gaps with synthetic data. Synthetic data here means taking recordings of real work and varying lighting, angle, object position, and work environment to mass-produce virtual training data. Where Figure adds people, the Korean plan means to reach part of the same variety by deforming one set of recordings.
Collection is not the only thing that splits; release splits too. The example the same article gives is Open X-Embodiment, built by Google DeepMind, Stanford University, and other research groups, which has put over one million real robot trajectories collected across many robot forms into a shared dataset. Figure built the opposite. As the announcement itself puts it, the pipeline is Figure-exclusive, which is why the pass rate and the rate per minute are both worked out inside it.
With factories, logistics centers, and farms as the stage, the problem from Section 4 is certainly lighter. No camera goes into somebody's living room, and consent to film gets organized at the level of a workplace. The same question stays, though. What standard separates the motion data worth training on from the rest, and whether that standard and the resulting pass rate will be reported as program outcomes. A national program finds budget and volume collected easy to announce first, and those two numbers are the same kind of number Figure announced.
This is why Pebblous keeps reading announcements like this one. Data work tends to get reported as volume collected, and the usable share of that volume stays inside. Collection speed is easy to announce and pass rate is hard. While the hard number goes unpublished, nobody outside can rank the competition, and inside, no evidence accumulates for comparing which collection design worked better.
An organization that collects data can turn three questions on itself while reading someone else's announcement.
- Can we say right now, as a single number, what share of the data we collected last quarter went into training?
- Does a record of why filtered data was filtered survive stage by stage, or does it disappear once dropped?
- Does our consent design distinguish the person who agreed to be recorded from the people who happened to be in the room?
Thank you for reading. The figures and quotations in this article were checked directly against Figure's Index announcement, the Index page, the privacy policy, and the Nscale partnership post, with the September figures cross-checked against Humanoids Daily and the Korean plan against the Korea Economic Daily report. If your organization measures the training adoption rate of the data it collects, we would be glad to hear what standard you use to draw the line.
Pebblous Data Communication Team
September 13, 2026
References
Figure AI Official Sources
- 1.Figure AI. (2026). "Introducing Index." 2026-08-25.
- 2.Figure AI. "Index."
- 3.Figure AI. (2026). "Privacy Policy." 2026-01-21.
- 4.Figure AI. (2026). "Figure and Nscale Sign Strategic Partnership." 2026-09-03.
News Coverage
- 5.Humanoids Daily. (2026). "Figure AI Reports Rapid Growth for Index, Surpassing 69,000 Weekly Active Users." 2026-09-08.
- 6.Korea Economic Daily. (2026). "Building a Robot Training Center to Collect Physical AI Data." 2026-06-25.