Executive Summary
A single story from Russell Brandom, AI editor at TechCrunch, published on September 18 and reworked two days later, reports an odd scene in the world model business. AMI Labs, founded by Yann LeCun, and World Labs, founded by Fei-Fei Li, have gathered plenty of attention and plenty of money, and the answers go vague at the question of who will buy the technology. This article looks at what that silence leaves behind, not for the companies building the models but for the company building the data.
The line that lasts longest in the story came from a supplier rather than a model company. Alex de Vigan, chief executive of Physicl, which sells 3D training data to companies in this field, said the company still does not know exactly what its customers are building. More information would mean more useful data. That reads differently from the familiar diagnosis that there is not enough data. The volume and the skill are both in place, and what is missing is the thing the data should match.
Sections 1 through 3 stay with what the story and the published material set out. Section 4 rereads the situation as a problem of data specifications, and that reading is this article's own rather than the story's.
Key Figures
Sources: company announcements and press coverage. Individual sources are linked in the body.
$1.03 billion
Seed round AMI Labs took
Announced in March 2026 at a pre-money valuation of $3.5 billion, the largest seed round on record out of Europe
3 to 5 years
Horizon LeCun gave for general systems
LeCun told AFP that talks with corporate partners start within one to two years, and that fairly universal intelligent systems are the goal in three to five
$1.23 billion
Total raised by World Labs
Another billion dollars arrived in February 2026 alongside the full launch of Marble, the first product. Autodesk anchored the round with $200 million
Millions
Simulation-ready 3D assets Physicl opened with
Announced at NVIDIA GTC in March 2026. Friction, mass and collision properties are derived from the geometry and materials automatically
Plenty of Money, No Named Product
Brandom wrote in the story about moderating the world model panel at the All In conference, an event unrelated to the podcast of the same name. The story names AMI Labs and World Labs as the two large players in the field, high on buzz and funding and fairly far down the list on any effort to earn revenue.
The story describes the core of a world model as handling spatial intelligence automatically. Its simplest form is a map drawn so that the world can be moved through, like the model that drives a self-driving car. And the same approach that gets a car through traffic can also carry a box across a room for a humanoid robot. So the available directions run wide, from robotics to controllable video to harder forms of autonomous driving.
Wide directions also mean no direction has been picked. Pressed on where commercial use would actually show up first, the answers got cloudy, Brandom wrote. Michael Rabbat, who leads world models at AMI Labs, answered on the panel: "We'll talk about it when we're ready to talk about it." The reply that came back by email was no different: "We're still in a research and building phase, so we're not talking publicly about any product plans or timeline."
None of which says nobody is doing anything. By the story's account AMI Labs already has a foot in manufacturing, biomedicine and robotics, and reaches AI software for doctors through a partnership. Several areas have been touched, and which one turns into a product first goes unsaid.
World Labs talks comparatively more. Brandom called this company's Marble probably the most fully developed product in the space, with demos ranging from straightforward media creation to explorable environments for video games to CGI effects. Marble came out in a limited beta in November 2025 and launched fully in February 2026. That same month the company raised another billion dollars. Autodesk anchored the round with $200 million, and Nvidia and AMD joined as well. Total funding reached $1.23 billion. What comes after this product is blank all the same.
Why No One Speaks First
Two strands run through the explanation the story offers. One is competition and one is money.
The competitive logic is plain. Naming the use for a technology exposes the road to market, and once that road is visible a rival can raise money to travel it too. Brandom put it this way: "Cixin Liu fans will recognize this as a dark forest scenario: If you don't know who else is in the woods, it's best not to attract attention." Deep funding makes the calculation stronger rather than weaker. The money that bought the quiet is feeding potential competitors at the same time.
The money question is how long a company can afford to stay quiet. AMI Labs was set up in Paris in December 2025 and took $1.03 billion as a seed round three months later, in March 2026. The pre-money valuation was $3.5 billion, the largest seed round on record out of Europe. Chief executive Alexandre Lebrun told TechCrunch that this is "not your typical applied AI startup that can release a product in three months, have revenue in six months," describing a plan that starts from fundamental research. LeCun told AFP that talks with corporate partners would open within one to two years and that the goal within three to five years is to produce "fairly universal intelligent systems." Press accounts of the round describe a company expected to stay a research organization with nothing to sell for roughly five years.
One more thing sits on top of that. The phrase "world model" is itself still broad. Lebrun once predicted that "in six months, every company will call itself a world model to raise funding." A broad label lets an investment story hold together without specifying any use. The cost of keeping quiet drops by that much.
World Labs, for its part, has said a fair amount. On June 3, 2026 the company published a taxonomy that splits the term three ways. A renderer "outputs observations in the form of pixels meant for human eyes." A simulator "outputs state: a geometrically, physically or dynamically faithful representation of the world that humans and computer programs can both compute on and interact with." A planner "outputs actions." Of the three, the piece judged that "the simulator gets the least public attention, and is the most consequential of the three." Data comes up too. Three-dimensional data with explicit geometry, material properties and physical annotations is "orders of magnitude scarcer than the internet video that renderers train on," the piece noted.
So even the shortage has been published. That still does not answer the question the company making the data is holding. What is scarce has been named, and how accurate that data has to be on which task before it earns its price remains open. A list arrived, not an assignment.
Put competition and funding together, the two strands above, and the silence at these companies is reasonable enough. Little is lost by saying nothing and much is lost by speaking. The arithmetic only balances inside the company boundary, though. Outside it are people whose work has to match whatever that company is building.
The Cost Falls on the Side Making the Data
Brandom met Alex de Vigan, chief executive of Physicl, in a hallway at the same conference. Physicl is a French company that left stealth at NVIDIA GTC in March 2026, launched by members of the Nfinite team, who had handled 3D data for retail. At the starting line it announced a library of millions of 3D assets and environments ready to drop into a simulator, and pointed at three areas: robotics, world models, and models that handle images and language together. The same announcement lists Meta, DeepMind, World Labs and Getty Images among the teams the platform already supports. On its own site the company describes how those assets are built: raw inputs become physics-tagged 3D with "geometry cleaned, materials resolved, friction, mass, and collision properties derived automatically."
The founding logic de Vigan set down in the announcement is clear. "Every major advance in AI has required a new data layer," and for physical AI the missing layer is "structured, spatially consistent, physics-aware data that models can actually learn from." A condition rides along at the end of that sentence. Existing is not enough; the form has to be one a model can take in.
And a company like that is working without knowing what its customers are building. That its data has been useful is known. What it was useful for is not. So de Vigan's remark reads as a statement about specifications rather than a complaint: "I wish they would tell us more. We could build more useful data if we knew what they were working on." The ability to build is blocked by an absence of purpose.
The difference shows up in a single asset. Say a warehouse bay gets built in 3D. If a robot arm will learn to pick things up there, the friction on the surfaces a hand touches, the mass of each object and the small deformation that appears under a grip all have to be right. If video or a game environment will be built from it, appearance and lighting and the variety of the scene come first. The cost multiplies several times over if both are matched at the highest level, and with neither one specified the asset falls a little short on both sides.
A direction to improve in goes missing along with the purpose. Knowing which assets were the good ones is what lets the next batch be pushed that way, and without knowing which task those assets went to, no standard separates good work from wasted work. Volume grows and quality scatters in every direction.
Here is where this parts from the existing discussion. Talk about the data bottleneck in physical AI has mostly settled on volume and cost. Demonstration data for a robot to learn from has to be gathered by a person moving a robot, so it cannot be scraped the way web documents can, and scale never arrives. What de Vigan points at sits next door. The capacity to build is already standing, and there is no target to aim at.
Solid evidence backs the volume diagnosis too. In a post from September 2025, the data company Scale AI reported that the relevant open datasets available at the time, including DROID and Open X-Embodiment, "offer only about 5,000 hours of interaction data combined." The same post counts "more than 100,000 production hours completed at our prototyping laboratory in San Francisco" to fill that gap. The shortage is real and closing it costs a great deal. Physicl too, in its announcement, measured itself against neighbors: it supplies simulation-ready 3D data "much like Getty Images serves licensed visual content or Scale AI provides labeled training data." The gap de Vigan pointed at stays exactly where it was even after the volume is filled.
An objection is available here. World models are built to serve many uses in the first place, so why not build data that serves many uses as well? Fair enough, and it moves the problem rather than removing it. Serving many uses is also a specification, and somebody has to write down the range. How many kinds of object to include, how precisely to match a friction value: budget always runs out somewhere, and where the range goes unstated, the guesswork of whoever is building fills the gap. With no way to check later whether it was right, a gap filled by guesswork hardens into place.
Why Pebblous Is Watching This Story
One distinction stays with us whenever we talk about AI-Ready Data. Having data and having data prepared for training are not the same thing. This story touches the step just before that distinction. No one can call data prepared without first knowing what it will be used for. When the use is left blank, the state called ready loses the ground it would be defined on.
The same question, turned toward our own work, produces a spec sheet. When data gets built by an outside supplier, the document handed over usually spells out form. File standards, label schema, class definitions, review criteria, delivery dates. Purpose, by contrast, rarely appears. What the data is meant to help judge, which situations a wrong model answer would hurt in, how good work will be confirmed: none of it is in the document. A supplier who receives that kind of spec sheet stands where the chief executive of Physicl is standing.
A spec sheet with the purpose left out gives up the ground for judgment. Meeting a case the rules do not cover, a supplier has nothing to lean on. An arbitrary call fails the review, and a question asked every time pushes the schedule back.
Language for conveying purpose is already in use across the industry. Scale AI has said that its robot data carries more than motion, layering "semantic detail that encodes intent, task structure, and failure modes." Writing down what counts as failure requires settling first on where the data will be used. The intent named there belongs to the person demonstrating, though, not to the party buying the data. The boxes a supplier can fill in alone and the boxes only the commissioning side can fill stay separate to the end.
This does not end with an instruction to write the purpose down in full. Writing a purpose out in detail also sends a company's plans outside the building. That is exactly the reason the world model companies keep quiet. So the question actually worth solving is not whether to disclose, but whether the place a dataset should aim at can be conveyed without exposing the plan. One route hides the final use and agrees on the evaluation method up front. Another hands over a handful of examples of cases where an error would hurt. Four checks give a rough position for a spec sheet as it stands today.
- Does page one of the spec sheet say, in a single sentence, what this data is meant to help judge? With only the formal rules and no such sentence, a supplier matches the surface.
- For cases the rules do not cover, does the sheet include adjudicated examples a supplier can work from? Ten ambiguous cases beat ten pages of taxonomy.
- Does the supplier know before starting how we will evaluate the delivery? Keeping the scoring sheet hidden until the end means the next round will not improve.
- If the purpose cannot be disclosed in full, has the part that may stay hidden been separated from the part that must be passed along? Hiding all of it is not security but surrender.
When the world model companies will speak is not ours to know. The scene this story leaves behind will stay useful longer than their roadmaps. Where purpose does not flow between the side making data and the side using it, the document joining them can run to any length and still carry only form. The problem does not come from a shortage of money, so money will not resolve it.
Thank you for reading this far. The facts this article cites can be checked by anyone in the TechCrunch story and the Physicl announcement. We would be glad to hear whether the last data spec sheet you sent outside carried its purpose, and if it did not, what stood in the way.
References
Primary Reporting
- 1.Brandom, R. (2026). "World model companies are keeping a lot of secrets." TechCrunch.
- 2.TechCrunch. (2026). "Yann LeCun's AMI Labs raises $1.03 billion to build world models." TechCrunch.
- 3.The AI Insider. (2026). "Fei-Fei Li's World Labs Raises $1B in Fresh Funding to Advance Development of World Models." The AI Insider.
Industry & Official Sources
- 4.Li, F. (2026). "A Functional Taxonomy of World Models." World Labs (Substack).
- 5.Nfinite / Physicl. (2026). "Physicl Launches the Data Infrastructure Layer for Physical AI at NVIDIA GTC." PR Newswire.
- 6.Physicl. (2026). "Physicl — Physical AI Data Infrastructure." physicl.ai.
- 7.Scale AI. (2025). "Physical AI." Scale AI Blog.