Executive Summary
This report maps the world-model landscape for Physical AI as of September 2026 and asks what synthetic data actually buys. The short answer: synthetic data does not lift a robot's ability evenly. It fills the specific weakness the data was aimed at. Feed in trajectories built by randomizing object placement, and the success rate jumps in the condition where placement changed, while the condition with distractor objects nearby and the condition with altered lighting barely move. Add a second batch that varies the environment instead, and now the distractor cells climb steeply while the placement cell inches forward. Meanwhile the score measured inside the training distribution hardly moves across both rounds of augmentation. What synthetic data bought was not skill. It was tolerance for changed conditions.
The basis for that reading is not a paper's headline number but the condition-by-condition table in its appendix. The authors ran four tasks under five conditions, hundreds of trials per data regime, and recorded twenty cells. The body of the paper reports three averages. An average alone cannot answer why the number rose. The same shape repeats on the metrics side. A world model built for robot policy evaluation reports a large jump in its average score, but split into its six metrics, only two rose, image quality actually fell, and trajectory accuracy, the metric closest to action, sits lowest of the six. A position paper on world-model evaluation gave this pattern a name, a mismatch between claims and evidence, then laid out eight levels of evidence and reported that the middle levels are empty. The most striking gap: almost nobody measures whether a policy's improved score came from exploiting holes in the generative model.
So this report does not rank models. It hands over the questions worth asking instead. In which condition did the number rise, and by how much. Was it compared against real data collection at a matched cost. Which cell moved, rather than what the average did. The stakes are concrete: a government-funded Korean world-model program, running on a two-year budget, has set its core goal as raising a real robot's task success rate by at least 20 percentage points over a no-world-model baseline. A gain reported without fixing the condition cannot be graded. What remains is the ordinary work of diagnosing data at the level of the task episode, naming the conditions to be filled, and measuring again under the same conditions.
+7.5%p
Gain inside the training distribution
Across both rounds, after 130 synthetic trajectories were added. Over the same span, the unseen-object condition rose 40.0 points
+33.7%p
Distractor condition, second round of augmentation
The same condition gained only 5.0 points in the first round. Change what the data aims at and the cell that moves changes too
52.5%
Final success rate under changed lighting
Still the lowest cell in all three data regimes after two rounds. Two-thirds of the same regime's 80.0% in-distribution score
0.3561
Trajectory accuracy behind an average of 0.6834
Three other metrics in the section 5 breakdown sit at 0.88–0.93. The cell closest to action is under half the average
One name, six different things
Start with an expiry date. Every release status and spec below is a snapshot taken in September 2026. In this field model cards are revised monthly, and a model listed as "coming soon" in May shows up in July with weights attached. The tables here are material for a judgment, not the judgment itself. When the purchase decision actually arrives, open the same pages again.
Before any of that, the vocabulary needs sorting out. In conference slides and vendor one-pagers, the phrase "world model" currently points at several different objects, and those objects promise different things. When one name covers all of them, a buyer cannot tell what is being bought.
1.1Start with a robot moving a cup
Picture the simplest possible scene. A robot arm slides a cup across a table. Now feed three different actions into a model that predicts what happens next: nudge the cup gently, shove it hard, do nothing at all. A good model should produce three different futures. The gentle nudge moves the cup a little, the hard shove tips it over or sends it off the table, and doing nothing leaves the cup where it was.
That property, where changing the action changes the predicted future accordingly, is called action-conditioned prediction, and it is the minimum requirement for a world model used in robotics. A model that returns a plausible-looking video no matter which action you feed it is a video generator, not an environment model for a robot. One metric later in this report makes the distinction painfully clear: a completely static video, with nothing moving at all, scores at the top on certain video-quality measures.
Action-conditioned prediction, illustrated (original Pebblous diagram). Feed different actions into the same current scene, and the predicted next scene should change accordingly: no action leaves the cup where it was, a gentle nudge moves it a little, and a hard shove sends it off the table. A model that returns the same scene regardless of the action fails action-conditioned prediction.
1.2The question that separates five concepts
Five terms get used interchangeably on the ground: world model, VLA, physics simulator, digital twin, synthetic data. Memorizing five definitions is less useful than holding one line per term about what to check before adopting it. The third column of the table below is that question.
| Concept | Primary role | Question to ask before adopting |
|---|---|---|
| World model | Learns environment structure or dynamics to represent, generate, predict | What does it take as input, and which future does it predict? |
| VLA | Takes vision and language, outputs actions | Does it fit our robot's action representation and control stack? |
| Physics simulator | Computes state changes from specified physics models and parameters | Are the contact, friction, and sensor models valid for our task? |
| Digital twin | A virtual model tied to one specific real asset | Which measurements is it synced to, and what is the error? |
| Synthetic data | Data produced by models, simulation, or transformation | For which training or evaluation purpose, and how is it screened? |
Five concepts, their roles, and the question each one raises. The boundaries are softer than they look. A VLA can carry prediction inside it, and a simulator can be combined with a learned world model. Not every 3D scene is a digital twin tied to a real installation.
1.3Six objects under one name
The practitioner's instinct here, ask about inputs and outputs rather than the name, has an academic counterpart that arrived at the same place. A position paper posted in June 2026 by researchers at the State Key Laboratory for Novel Software Technology at Nanjing University counts the objects the word points at, right at the top of its abstract.
"The term now refers to several different objects: action-conditioned environment models, latent imagination models, future-video predictors, interactive neural simulators, latent predictive representations, and synthetic-data engines."
Six objects, then: a model of the environment that takes actions as a condition, a model that imagines in latent space, a predictor of future video, an interactive neural simulator, a latent predictive representation, and an engine for producing synthetic data. The last of the six governs the back half of this report. A world model is not only the robot's sparring partner during learning. It is also being used as a factory that stamps out robot training data.
1.4Different claims need different evidence
The same paper goes one step further. These six objects make six different claims: that the model generates good future video, that it can evaluate a policy, that it can improve a policy, that it supports planning and control, that it can supply synthetic data for training, and that it has learned a useful representation. What the paper guards against hardest is treating those claims as one. "These claims are not interchangeable. Evidence for one of them is not automatically evidence for another." Evidence for one claim does not carry over to the next.
So "world model inside" tells a buyer almost nothing about what is for sale. A document that argues it can improve a policy on the strength of good-looking video is worth something quite different from one that argues the same thing on the strength of correctly ranked policies. Section 5 lays out, claim by claim, which evidence each of the six actually requires.
The academic lineage and taxonomy of world-model research is covered separately in our world model survey. This report does not repeat that taxonomy; it picks up where the taxonomy turns into a practical decision.
September 2026: what is open, what is closed
This section is not a leaderboard. The world-model families released so far run on different robots, train on different data, and are scored on different evaluation tasks. A table that lines up their success rates side by side would mislead whoever reads it. The axis compared here is therefore not performance but openness: what is released, what is held back, and how far each team documents its own limits.
The openness spectrum of world-model families (original Pebblous diagram). What this section compares is not robot performance but openness: how far outward the model's claims can be checked. The further right, the more of the claim an outside reader can verify directly.
2.1Where seven families stand
The table below covers the seven families whose public documentation could be verified as of September 2026. The middle column says what each family produces; the right column says how far an outsider can check that claim.
| Family | Core function | Release status as verified |
|---|---|---|
| Cosmos 3 NVIDIA | Action-conditioned video and action generation, tuned for robots and vehicles. Three sizes: Super, Nano, Edge | Weights released for all three sizes (Hugging Face); model cards include a limitations section |
| Genie 3 Google DeepMind | Generates explorable virtual worlds from text or images, playable in real time | No weights, no API. Reachable only through a top-tier paid consumer subscription (Project Genie) |
| Atlas World Labs | Reconstructs real spaces from images and video, outputs novel-view frames and explicit 3D | Partner-only early access. No weights, no public API, no pricing, no performance figures |
| Marble World Labs | Generates explorable 3D scenes, exportable as splats, meshes, or panoramas | Cloud API billed in credits. Weights closed |
| Dreamer 4 Published September 2025 | Trains agents inside a latent-space world model. Real time on a single GPU | No official weight or code release from the authors. Only unofficial reimplementations are public |
| GigaWorld-1 GigaAI and Tsinghua University | World model built for robot policy evaluation. Two sizes: Nano and Plus | Code, checkpoints, curated dataset, and toolkit all under Apache-2.0 |
| V-JEPA 2.1 Meta AI | Latent predictive representation learning; plans manipulation via model-predictive control | Code and checkpoints released (ViT-B/L/g/G) |
Core function and release status for seven families, as verified in September 2026. Dreamer 4 is the only entry here that is a year old, so its publication date is listed alongside. Results obtained in Minecraft should not be read across to physical robot performance. The wider industrial picture is covered separately in our Physical AI industry landscape.
2.2The trap in a side-by-side spec sheet
Spec comparisons are useful, with one trap built in. Some families output video; others output not video but a 3D scene you can walk into. Frames per second is not a meaningful axis for the second group at all. Leave that difference unmarked and a reader will conclude that the families producing 3D output are simply behind on frame rate.
| Family | Max generation length and resolution | Action-conditioned | License and access |
|---|---|---|---|
| Cosmos 3 Super (64B) | 5–400 frames, input up to 720p | Yes | Weights released, NVIDIA Open Model License |
| Cosmos 3 Nano (16B) | 5–400 frames, 720p / 480p / 256p | Yes | Weights released |
| Cosmos 3 Edge (4B) | 50–150 frames, 640×360 for robot control | Yes, limited to a supported list | Weights released; model card states OpenMDW1.1 |
| Genie 3 | 60 seconds, 720p, 24fps | No (explore and remix) | Closed, paid subscription product |
| Atlas | 1 minute, 1440p* | Partial (observations generated along a path) | Partner-only early access |
| Marble | Not applicable* | No (camera paths only) | Credit-billed API |
| Dreamer 4 (2B) | Continuous real time, 21fps on a single GPU | Yes (keyboard and mouse) | No official release |
| GigaWorld-1 (1.3B / 5B) | Long-horizon rollouts demonstrated to 40 seconds | Yes (control injection) | Fully open under Apache-2.0 |
* Atlas and Marble produce explorable 3D output rather than video, so frames per second and generation length are a different kind of axis for them and cannot be compared line by line with the rest. Even inside the Cosmos 3 family the license string splits: documentation for Super and Nano says NVIDIA Open Model License while the Edge model card says OpenMDW1.1, and the public pages alone were not enough to confirm whether those are two names for the same license. Secondary summaries report a different Edge parameter count from the Cosmos 3 technical report, so this report uses only the model card value of 4B.
2.3Is your robot on the list?
A "yes" in the action-conditioned column does not mean the model will accept your robot's actions. To take actions as a condition, a model has to know that robot's action dimensionality and normalization scheme. The model card for the edge-sized model names the action formats and dimensions it supports, which turns an abstract warning into a checklist you can verify.
| Action format | Dim | Action format | Dim |
|---|---|---|---|
| Generic camera motion | 9D | Agibot | 29D |
| Autonomous vehicle | 9D | UR | 10D |
| Egocentric navigation | 57D | Google robot | 10D |
| Franka Panda, single arm | 10D | WidowX 250 | 10D |
| Franka Panda, dual arm | 20D | UMI | 10D |
Action formats and dimensions listed on the Cosmos3-Edge model card. Action input arrives as JSON holding a two-dimensional array of frames by dimensions.
How to read it is simple. If your robot is not on this list, the model is not usable as-is, and a post-training setup that defines the action space, dimensionality, camera placement, and normalization comes first. It is worth reading the list for what is absent as well. Quadruped locomotion and whole-body humanoid control do not appear on it.
2.4Teams differ in how they write down their limits
The most valuable finding in this section is not a spec number but a difference in how the documents behave. The teams that opened their weights also named what their model is merely approximating. The limitations section of the edge model card opens like this.
"Because the model lacks an explicit physics simulator, 3D geometry, 4D space-time evolution, object permanence, contact dynamics, and physical laws are only approximated — producing artifacts such as disappearing or morphing objects, unrealistic collisions, and physically implausible motions."
Without an explicit physics engine, the card says, 3D geometry, space-time evolution, object permanence, contact dynamics, and physical law are all approximations, and the visible consequences are objects that vanish or morph and collisions that do not behave. The same document adds temporal inconsistency and drift in action and state over long horizons, then states flatly that the output must not be treated as physically accurate simulation or as safety-certified decision making. From a buyer's seat this is good news rather than bad. The document tells you what to verify.
The contrast sharpens on the other side. On the official blog of a team that kept its weights closed, there was no limitations passage at all. The pages describe reproducing the physical interaction of rigid, articulated, and deformable bodies, yet no statement about the accuracy of that physics could be found. The robotics explanation reduces to one sentence: "As a simulated robot moves through space, Atlas also generates the RGB and depth data its sensors would observe along the way." What it generates is observation. Contact forces and joint torques are not on the list.
One qualification is owed here. Open weights by themselves do not lower the cost of adoption. Inference hardware, post-training, robot data conversion, and screening costs all have to enter the same comparison. Some names are also missing from the table on purpose. For several newer research outfits, cited both in Korea and abroad, no open weights, public API, or comparable performance table could be confirmed as of the research date. Unverified is not the same as nonexistent, so nothing is declared absent, but nothing unverified was placed in the table either.
Openness also makes headlines checkable, which the latent-representation side demonstrates nicely. The abstract of one open model reports that grasping success on a real robot rose 20 points over the previous version. Open the tables and that figure comes from one of three tasks, grasping, while reaching and pick-and-place are unchanged. The 20 points do not come from a single inference setting either: at a one-step horizon the gain is 10 points, and reaching 20 requires changing the horizon, the number of repetitions, and the sample count together. The authors write themselves that the remaining failures come from gripper action planning rather than spatial understanding. The representation improved; the bottleneck stayed on the action side.
Both the ingredient and the product
Synthetic data and world models point at each other in two directions, and separating those directions keeps the rest of this report from tangling. In one direction, synthetic data trains the world model: simulation and transformation fill the gaps where real robot data is thin, and a predictive model grows from that. In the other, the world model produces synthetic data: a trained model stamps out new scenes and new trajectories, and the output becomes training material for a robot policy. The same word, pointing at an ingredient in one sentence and at a product in the next, changes what has to be verified.
Synthetic data as ingredient and as product (original Pebblous diagram). The same word names the material that trains a world model at the top and the output a trained world model stamps out at the bottom. The dashed loop marks that the same word changes what has to be verified depending on which role it is playing.
3.1Purpose decides the verification
Why you are using synthetic data determines what you have to check. The table below takes five representative purposes and pairs each with the data it needs and the verification that decides whether that data is any good. That verification column is the most operational thing in this report.
| Purpose | Data required | Representative verification |
|---|---|---|
| Learning to predict the environment | Links between current observation, action, and future observation | When the action changes, does the outcome change plausibly? |
| Imitation learning of behavior | Observations, instructions, executable action trajectories | Do observation and action match, and is the target task completed? |
| Improving perceptual robustness | Observations with altered appearance conditions and valid labels | Do object and label stay aligned after the background changes? |
| Pre-deployment policy evaluation | An environment that responds to the policy, plus success judgments | Does the policy ranking in virtual evaluation match reality? |
| Learning failure and recovery | Failure, abort, and recovery segments with recorded causes | Are successful and failed behaviors separated for the training purpose? |
Purpose, required data, and the representative verification for each. The two verifications set in bold return in section 5. The earlier one is the point the academic ladder of evidence calls the intervention threshold; the later sits two rungs above it, at policy-ranking agreement.
A trap hides in the last row. A model that predicts outcomes can learn the conditions under which failure occurs, so failure segments are an asset there. Mix those same failed trajectories into a policy that imitates successful behavior, under the same goal label, and the policy learns behavior nobody wants. The identical data is raw material on one side and contamination on the other. Failure, abort, and recovery segments therefore have to be marked separately, and a synthetic data package that arrives without those marks can serve only one of the two uses.
3.2The problem each route leaves behind
There is more than one road to synthetic data for robot learning. Six routes are in active use, and because each yields something different, each leaves a different problem behind. Read the table from its right-hand column. What remains unresolved decides an adoption call more often than what you gain.
| Generation route | What you get | What still needs verifying |
|---|---|---|
| Physics simulation with domain randomization | Observations, states, actions, sensor labels | Gap between the parameter range and real dynamics |
| Demonstration retargeting and replay | Action trajectories executed in new layouts | Reachability, contact, whether the task actually succeeds |
| Structure-conditioned video translation | Appearance variants attached to existing trajectories | Object boundaries, contact timing, occlusion versus labels |
| Generated video plus inverse dynamics | Video and inferred actions | Action inference error, robot compatibility, executability |
| Action-conditioned world model rollouts | Future observations under the input actions | Long-horizon error accumulation; policies exploiting model error |
| Reconstruction and simulation of real scenes | Virtual scenes matched to a real site | Unobserved regions, physical properties, measurement error |
Six generation routes and the verification work each leaves behind. The right-hand cell of the fifth row reads policies exploiting model error. Section 5 shows that almost nobody measures that cell.
The representative systems on each route are mostly familiar by name. Demonstration retargeting multiplies trajectory count; video translation widens visual variety; dream-generation approaches infer actions back out of generated video. One caution: results from driving video should not be generalized to assembly or deformable-object manipulation. The concrete shape of a pipeline is covered in our synthetic data pipeline piece and is not repeated here, and techniques for physically verifying generated output after the fact are collected in a separate article.
3.3A trajectory marked successful is not a clean trajectory
The argument that generated data needs a screening layer is not new. A 2024 paper introducing a large-scale robot simulation dataset wrote it into its own limitations section. That it comes from the people who built the synthetic data is what makes it worth quoting.
"While the generated trajectories are technically considered successful, many exhibited undesirable effects, such as jerky motions and collisions. Many of these behaviors can be automatically detected by checking simulation states and trajectories exhibiting such behaviors can be discarded."
Three things sit in that one paragraph. First, synthetic trajectories carry bad motion even after being scored as successes. Second, success rate is therefore not a data quality metric. Third, those behaviors can still be caught automatically by inspecting simulation state.
That last point is the one that matters. It is not a complaint; it records that the problem is solvable. When a synthetic data package arrives, the question is not "what is the success rate" but "among the trajectories scored as successful, how did you filter out the ones with jerky motion and collisions?" Without that procedure, a dataset with a 100% success rate can teach a robot bad habits.
From 40.8% to 69.0%: which cells moved
This is the center of the report. A paper released in August 2026 fed world-model-generated synthetic trajectories into robot policy training and reported the change in success rate. Overall success went from 40.8% to 69.0%, a gain of 28.2 percentage points. Copying that number across would not justify an article. The value is in the condition-by-condition table in the paper's appendix. The authors measured twenty cells and reported three averages in the body.
4.1What was measured, and how
The numbers only read correctly after the experimental design. There are four tasks: moving a bowl, folding a towel, placing a cup, opening a drawer. Each task starts from 25 real robot demonstrations, and synthetic trajectories are added in two stages. The first stage adds 65 trajectories built by randomizing object placement; the second adds another 65 that vary the environment. Training data therefore comes in three regimes per task: 25, 90, and 155 trajectories.
One control matters a great deal. Policies in all three regimes start from the same pretrained checkpoint and share the same training configuration. In the authors' words: "All policies are initialized from the same π0 checkpoint and use identical training configurations; only the composition of the training data differs across regimes." Nothing but the data was changed, and they say so explicitly.
Evaluation ran on a real robot under five conditions: in-distribution, changed object placement, added distractor objects, unseen object instances, and changed lighting. Every task-condition pair was tested 20 times, so four tasks times five conditions times twenty trials gives 400 trials per regime, or 1,200 physical trials across the three regimes. A condition-level table at that scale living only in an appendix is the surprising part.
4.2Unfold the twenty cells
The figure below places the three regimes side by side within each of the five conditions. Light gray is the baseline trained on real data only, dark gray is the first round of augmentation, orange is the second. Look first at the spacing between the three bars: it draws a completely different shape in each condition.
Real-robot success rates for three training data regimes across five evaluation conditions (author-reported, appendix table). Values are averaged over four tasks, 400 trials per regime. The argument of this section is that the shape of the three bars differs from condition to condition.
4.3The same numbers, written as increments
Rewriting the same values as increments makes what happened legible at a glance. The table below converts the gaps between the bars above into numbers. Cells in orange are where a given round of augmentation moved things sharply; cells in gray barely moved at all.
| Evaluation condition | Real only | Round 1 | Round 2 | Total |
|---|---|---|---|---|
| In-distribution | 72.5 | +3.8 | +3.7 | +7.5 |
| Changed object placement | 37.5 | +25.0 | +7.5 | +32.5 |
| Added distractor objects | 33.8 | +5.0 | +33.7 | +38.7 |
| Unseen object instances | 30.0 | +12.5 | +27.5 | +40.0 |
| Changed lighting | 30.0 | +5.0 | +17.5 | +22.5 |
| Overall average | 40.8 | +10.2 | +18.0 | +28.2 |
Increments by condition, in percentage points, obtained by plain subtraction across the three regime columns of the appendix table. The overall-average row is the set of three numbers that appears in the paper's body; the five rows above it lived only in the appendix.
4.4What goes in decides what goes up
First, only the condition the augmentation aimed at moves much. Adding trajectories with randomized object placement lifted the placement condition by 25.0 points, while distractors and lighting each crept up 5.0 points. Adding environment-varied trajectories next moved distractors by 33.7 points and unseen objects by 27.5, while placement gained only 7.5. What you put in decides which cell rises. The effect of synthetic data is a function of dataset composition, not dataset size.
Second, inside the training distribution almost nothing happens. The in-distribution score went from 72.5% to 80.0%, a total of 7.5 points across both rounds. Over the same span, the unseen-object condition gained 40.0 points, more than five times as much. What synthetic data bought was not more skill at the task itself but the ability to hold up when conditions change. That is not bad news; robustness is usually where deployments break. It does mean that "the success rate went up" and "the robot got better at the task" are not interchangeable sentences.
Third, lighting stays last throughout. The lighting condition is the lowest cell in all three regimes, and even in the final doubly-augmented regime it reaches only 52.5%, two-thirds of that regime's 80.0% in-distribution score. It did move both times, 5.0 points in the first round and 17.5 in the second, yet after the other conditions climbed into the seventies, lighting remained near half. That appearance-varied data went in and lighting did not follow is itself diagnostic information. The variety that was added may not have covered the lighting changes a real site produces.
Go one level deeper, into individual tasks, and non-monotonicity shows up too. In-distribution success on towel folding runs 65%, 70%, 65% across the three regimes. The final regime scores lower than the middle one. The task-level aggregate (32%, 42%, 59%) hides that cell entirely. The counterexample to "adding synthetic data raises every cell" sits inside the same table.
4.5What this table does not answer
The conditionals have to be attached. Every number above is author-reported, obtained on four tasks with a single robot. It is not a Pebblous result, and it is not an independent third-party reproduction. One thing is also missing from the table: an arm that added the same number of real demonstrations. There is a group with 130 extra synthetic trajectories and no group with 130 extra real ones. No sentence acknowledging that confound could be found in the full text either.
There is a second confound. Data quantity and data diversity changed together. As the set grew from 25 to 90 to 155, the range of conditions widened alongside it, so this table cannot separate whether the score rose because of volume or because of variety. That puts a ceiling on what the experiment can conclude. It shows that adding synthetic trajectories raises the condition those trajectories aimed at. It does not show that doing so beats real collection at the same cost.
4.6How would you grade a 20-point target?
Why this argument matters right now is easiest to see in a publicly funded program. On June 9, 2026, Korea's Ministry of Science and ICT and the Institute of Information and Communications Technology Planning and Evaluation held the kickoff meeting for a Physical AI leading-technology development program, at LG Science Park in Seoul. The program runs two years on a budget of about $23.4M (KRW 34B), is led by LG Electronics with ten industry and academic partners, and plans four rounds of iterative field validation over that period. The kickoff announcement states the core goal as follows (translated from the Korean).
The core goal is to maximize the world model's real-world simulation performance and its transfer to a robot foundation model, raising a real robot's final task success rate by at least 20 percentage points over a no-world-model baseline, exceeding the 14.5 percentage points cited as the current global best.
The 14.5-point figure quoted there as the comparison baseline needs care. This report could not locate the paper behind that number or the experimental conditions it came from. It is therefore not carried here as fact, only as something the program materials cite.
Yet a paper published two months after the target was announced already records 28.2 percentage points. Has the bar been cleared? Answering requires asking "under which condition." Inside that one paper the gain splits into 32.5 points for the placement condition and 22.5 for lighting, and by augmentation round it splits further, into 25.0 points and 5.0. A gain reported without fixing the condition cannot be graded. That is the operational conclusion of this report, and the next section shows the academic side reaching the same place by another road.
One detail in how the program divides the work stands out. One dedicated institute is assigned to building the data collection platform and gathering core data, and a second to data standardization and model validation, with a domestically built simulator as a separate track. Diagnosis and validation were carved out as roles of their own, which reads as a sign that the problem described in this section was recognized while the program was being designed. The three current limitations named by the lead organization's principal investigator at the kickoff point the same way: physical consistency, sustained continuous execution, and the need to retrain on data. Those are nearly the same items as the limitations list on the foreign vendor model card in the previous section. The fork in the road is not domestic versus imported. It is who measures the same limitations, and how.
The ladder of evidence is hollow in the middle
The academic side reached the problem from section 4 by a different route. The Nanjing University position paper quoted earlier diagnoses a recurring mismatch between claims and evidence as world-model evaluation metrics proliferate: papers routinely claim more usefulness than their own evaluation can support. To make that mismatch visible, the authors laid out eight levels of evidence.
5.1Eight levels, and the question each one asks
The eight levels are not a scoreboard. Each answers a different question, and the paper notes that the levels span several orthogonal axes, so a higher level does not mean a better model. What to read in the table below is not the level number but the question in the right-hand column. Its use is to check which of those questions the document a vendor handed you actually answered.
| Level | Name | The question that level asks |
|---|---|---|
| L0 | Surface plausibility | Does the output look like a realistic image or video? |
| L1 | Matching a recorded future | Does the predicted future match a held-out real trajectory? |
| L2 | Instruction and semantic agreement | Does the rollout match the instruction, the task, and the scene's meaning? |
| L3 | Physical plausibility | Does the rollout respect intuitive physical and geometric constraints? |
| L4 | Action controllability | When the action changes, does the task-relevant change actually follow? |
| L5 | Predicting reward and outcome | Does it predict success, reward, progress, or value accurately enough to decide on? |
| L6 | Policy ranking agreement | Does model-based evaluation agree with real or simulator performance? |
| L7 | Improvement in decisions | Does using this model actually make decisions better? |
Eight levels of evidence and the core question at each. The paper states that the levels are neither mutually exclusive nor monotone in practice: a video model can score high from L0 through L3 and still be weak at L6.
Two axes are enough to carry into practice. One is observation versus intervention. L0 through L3 inspect what the model produced; from L4 on, you change the input and check whether the result changes accordingly. The paper's name for L4 is precise: "the first level that clearly distinguishes a decision world model from a future-video prior." The verification set in bold in the section 3 table, whether the outcome changes plausibly when the action changes, is exactly this level.
The other axis is whether the policy is held fixed or optimized against the model. L6 uses the model to rank policies that already exist; L7 uses the model to make a policy better. That difference leads straight into the next subsection. Optimization means pushing against the model, and pushing is what exposes the model's soft spots.
5.2Only the top and bottom are crowded
The authors held these eight cells up against actual research. What they counted was not the whole survey but roughly 40 papers listed in two tables, the subset they describe as most likely to carry decision-relevant claims. The placement rule is stated openly: a paper is assigned to a level when it reports at least one quantitative metric aimed at that level's core question, regardless of how much emphasis it places on the result.
Three patterns come out. Two of them concern which cells are crowded and which are empty. First, L1 and L7 are the two most common cells, and a sizable share of the work reports only those two: how well the held-out trajectory was matched, and how much the final score rose. Second, in between them, L6, which tests ranking agreement for fixed policies, appears in only about a dozen papers, nearly all of them from the cluster devoted to policy evaluation. The bottom and the top of the ladder are crowded and the middle is empty. The remaining pattern carries different weight and gets its own subsection.
The eight levels of evidence and how often each is reported, drawn from the two patterns in the paragraph above. Solid rungs are the two most commonly reported levels, dashed rungs the three that are rarely reported, and the orange dashed line marks where observation gives way to intervention. The third pattern, printed in red on the L7 rung, comes in the next subsection.
What those two patterns leave behind, the authors write, is that the model's own contribution stays entangled with the optimizer, the reward model, and the data pipeline. The score went up, and what produced it is unknowable. Their caveat belongs here too: the count is "a reading of the metrics recorded in our tables rather than as a full audit of every paper's appendix," though they add that the qualitative conclusion does not hinge on small differences in the tally. So the cell counts in the figure should not be read as precise statistics. What to read is the shape.
5.3Nobody measures whether the model got gamed
That pattern is the single most valuable passage in this report. The paper calls it the most striking of its own findings.
"Most strikingly, we find essentially no work in these tables that reports an exploitability gap, a measured discrepancy between model-predicted value and real value for policies or action sequences optimized against the model … the specific failure mode that model-based optimization is known to provoke is, by the evidence of our own survey, almost never measured."
For policies or action sequences that were optimized against the model, essentially no paper in those tables measures and reports how far the model's predicted value drifts from the real one. The paper gives the reason as well: "planners and optimizers actively seek out regions where the model is overoptimistic, so that improving average accuracy does not prevent catastrophic failures at the optimizer's chosen points." Planners and optimizers go looking for the regions where the model is too optimistic, so raising average accuracy does nothing about the catastrophic failures waiting at the points the optimizer picked.
Almost nobody measures how much of a score gained from synthetic data came from exploiting holes in the generator. The policy seeks out the regions where the generator is optimistic. And in the section 3 table, the row for action-conditioned world model rollouts already lists "policies exploiting model error" as open verification work. The cell practitioners had written down as work to be done is the same cell the position paper's census confirms nobody measures.
5.4What evidence does a synthetic data claim require?
The same paper tabulates, for each of the six claims from section 1, which evidence is strong and which is weak. The row for the synthetic data claim is the spine of this report. The three cells are short, so the original wording is kept alongside.
| Cell | What the paper writes |
|---|---|
| Stronger evidence | "Controlled downstream learning gains under matched training budgets and learners" |
| Weaker evidence used instead | "Video aesthetics, FVD, or VLM preference over generated data" |
| Why it misleads | "Visually clean data need not contain the right task-relevant variation for learning." |
Evidence a synthetic data claim requires, and the evidence used in its place. In the policy-optimization row of the same table, the weak evidence is given as a single downstream success number without decomposition, and the reason as the fact that final success entangles the world model with the reward model, filters, optimizer, and data curation.
The first row meshes exactly with section 4. Strong evidence means a control arm with matched training budget and learner, and the experiment in section 4 has no such arm. This is not nitpicking from the outside: it is holding one side's own evidence standard up against the other side. Another row of the table points at the Korean program's target. If a single undecomposed success number counts as weak evidence, then a 20-point goal stated without conditions is exactly that shape.
What happens when the section 4 paper is placed on this ladder? The chronology has to be stated first. The position paper appeared in June and the section 4 paper in August, so the latter is not part of the position paper's census. The placement rule is public, though, which lets this report do the mapping directly. There are numbers for video prediction quality (L1) and final success rate (L7), and none for L4 through L6 in between. And there is one more step here. The decomposition that would supply the missing middle evidence was already sitting in that paper's appendix. The authors measured it and left it out of the headline.
5.5The average rises while image quality falls
The same shape shows up once more, this time in metrics. The world model for robot policy evaluation released by GigaAI and Tsinghua University in July 2026 reports a 14.9% gain in average score over a competing baseline on its own benchmark. The comparison has to be named precisely. That 14.9% is measured against the strongest general-purpose baseline (Wan 2.2 5B); against the strongest robot-specific baseline it is 11.6%. The two numbers must not be blended.
The paper also says what the average is made of: a normalized mean over six chosen metrics, namely aesthetic quality, image quality, JEPA similarity, semantic alignment, subject consistency, and trajectory accuracy. This report checked that the plain mean of the six values matches the average printed in the table to four decimal places. The weighting is therefore equal, and trajectory accuracy carries the same one-sixth weight as aesthetic quality. Now split that 14.9% by metric.
| Metric | General baseline | Robot-specific model | Increment | Share of the movement |
|---|---|---|---|---|
| Aesthetic quality | 0.3538 | 0.3534 | −0.0004 | −0.1% |
| Image quality | 0.6980 | 0.6765 | −0.0215 | −4.0% |
| JEPA similarity | 0.5853 | 0.9337 | +0.3484 | +65.5% |
| Semantic alignment | 0.8789 | 0.8926 | +0.0137 | +2.6% |
| Subject consistency | 0.8883 | 0.8883 | 0.0000 | 0.0% |
| Trajectory accuracy | 0.1643 | 0.3561 | +0.1918 | +36.1% |
| Mean of six metrics | 0.5948 | 0.6834 | +0.0886 | = +14.9% |
The 14.9% gain broken down by metric. Increments and shares were computed by this report from the values in the paper's own table. The paper's text confirms the column-to-value mapping directly: it reports achieving the best score on JEPA similarity at 0.9337, semantic alignment at 0.8926, and trajectory accuracy at 0.3561, and ties the best value on subject consistency.
The reading is the same as in section 4. First, two of the six cells are the whole story. JEPA similarity and trajectory accuracy account for 101.6% of the movement, and the remaining four sum to minus 1.6%. Aesthetic quality and image quality went down. The average rose while image quality fell. Second, the cell closest to action has the lowest absolute value. The average of 0.6834 is three cells in the 0.9 range hauling up one cell in the 0.35 range. It reads as 68 out of 100, while the action cell is 36.
Third, the cell the work aimed at did move a lot. General-purpose models score between 0.09 and 0.16 on trajectory accuracy, and the robot-specific model reaches the mid-0.3s. The paper's claim is true in that cell; what demonstrates it is that cell, not the average. Scaling the model 3.85 times raised the average from 0.6717 to 0.6834, while trajectory accuracy rose by 0.94%. Nearly four times the parameters mostly bought the ability to keep a scene consistent, and left the ability to follow actions precisely about where it was. Reading this as the paper inflating its numbers would be wrong. The paper splits the two comparison figures itself and offers the per-metric decomposition itself. What travels around unsplit is a one-line headline.
5.6A frozen video scores at the top
The more important passage in that same paper is where the team audited the metrics themselves. They measured how well each automatic metric correlates with actual policy-evaluation ability, and found metrics that correlate negatively. The sentence explaining why is the most valuable line quoted in this report.
"Surprisingly, we identify several metrics that correlate negatively with WMES, including Background Consistency (ρ = −0.45), Photometric Consistency (ρ = −0.42), and Interaction Quality (ρ = −0.11). The first two fail because a trivial baseline that generates a completely static video can achieve high appearance stability while ignoring all actions and therefore failing entirely at policy evaluation."
The first two metrics fail, in other words, because a trivial baseline that emits a completely static video can score high on appearance stability while ignoring every action. The easiest way to produce a video whose background never wobbles and whose brightness never shifts is to move nothing at all. And that video fails policy evaluation completely.
One correction has to be attached right here. Reading this as "appearance does not matter" gets it wrong. In the same study appearance splits in two. Video fidelity, which measures whether frames are sharp and coherent, correlated best at 0.78, while appearance stability, which measures whether a scene holds still, ran backwards at minus 0.44. The paper's headline conclusion was that what matters is not short-horizon visual realism but action-faithful consistency over long horizons, so quoting either side without the short-versus-long condition distorts both.
The third metric in the quotation is not a minor one either. Interaction quality is unreliable, the paper says, because the metric is obtained by asking a VLM about physical coherence, and today's VLMs do not judge physical realism robustly enough to separate fine-grained rankings. The position paper carries the same caveat: a VLM can reward output that merely looks plausible and can miss subtle action-relevant errors, so its judgments become interpretable when checked against executable outcomes, policy rankings, and real performance. Two independent sources point at the same place.
The scope condition is owed as well. The position paper states that its stance is neither against video metrics nor against VLM judging. If the intended use is future-video generation, video quality is a legitimate primary objective; the problem begins the moment evidence fitted to one claim is used to support a stronger one. It also insists that the mismatch is not hypothetical. Two separate benchmarks reported, independently, that perceptual rankings and functional rankings disagree and that video quality is not a reliable predictor of executability, which the paper describes as "published findings on real model sets, not a thought experiment."
5.7First place on a leaderboard, but which test?
Move the metrics discussion onto leaderboards and one practical question appears. A vendor announcement from May 2026 states that its model ranks first across nine benchmarks and leaderboards, grouped into world-generation accuracy, action policy, and vision. That sentence is vendor self-reporting with no absolute numbers attached, so this report does not carry the first-place verdict as fact. What it did instead was check who built each leaderboard and what each one measures.
- Of the nine, only RoboArena compares policies on real robot hardware. A multi-institution academic consortium runs pairwise policy comparisons on a shared robot platform. The rest are simulation, or they measure video and understanding: user-preference Elo, physical realism against real captured footage (Physics-IQ), and VLM-scored generation quality (PAI-Bench, R-Bench).
- Of the two grouped under "action policy," one (RoboLab) is a benchmark built by the announcing vendor itself, and it is measured in simulation.
- The two grouped under "vision" (VANTAGE-Bench and TAR) test fixed-camera and traffic video understanding, unrelated to robot manipulation, and their datasets are hosted on that vendor's account. Whether the operators are independent could not be confirmed either way, so only the dataset ownership is stated here.
Publishing your own benchmark is ordinary practice and nothing to sneer at. The value is that a buyer gains items to check. Who built this leaderboard, and what does it measure?
Lay those three items over the earlier metrics discussion and section 5 reaches its conclusion. A ranking on a leaderboard that scores physical plausibility by asking a VLM tells you almost nothing about what a robot can actually do. Two independent sources hold that sentence up.
What about using a world model as an evaluation tool, then? A study led by Stanford researchers used an action-conditioned video generation model as a stand-in for the real environment and evaluated policies inside it, reporting that in-model success rates correlate highly with real ones and that relative rankings are preserved across policy versions, sizes, and checkpoints. The same study records its limits: robot motion is reproduced faithfully, while generating realistic object interaction remains hard. For separating policies by rank it is usable today, and on tasks where contact decides the outcome, physical verification still stands.
Diagnosing data one episode at a time
This section turns the criteria built in the previous two into a working procedure. It makes no new claims; it translates what sections 4 and 5 established into what to measure and what to record.
It starts by changing the unit. The habit of measuring data quality per image transfers badly to robot data. The problem is not one blurry frame; it is that sensor timestamps were misaligned during the 0.3 seconds when the hand met the object. The atomic unit of robot data is not the frame but the task episode, a stretch of time in which observation, action, and outcome are bound together, and quality can only be judged inside that stretch. The layered structure of robot learning data is covered separately in our robot learning data layers piece.
6.1The difference between a diagnosis and a complaint
The eight items below are the lenses to use when inspecting data episode by episode. A table that stops at finding problems is a list of complaints, not a diagnosis; it earns its keep only when the right-hand column names the improvement each item leads to.
| Diagnostic lens | Example of what to measure | How it connects to improvement |
|---|---|---|
| Completeness | Missing stretches of observation, action, or state | Recover the required modality or restrict the scope of use |
| Temporal alignment | Offsets between sensors and action latency | Correct the timing, then re-check cause and effect ordering |
| Semantics and units | Frames, rotation representations, units, instruction agreement | Standard conversions with a recorded conversion history |
| Distribution by condition | Combinations of task, object, lighting, friction | Choose augmentation conditions from real exposure and real failures |
| Duplication and skew | Repetition of near-identical frames, trajectories, scenes | Select by event and preserve rare but important stretches |
| Physical validity | Interpenetration, implausible contact, discontinuous motion | Physics checks plus sample verification by real execution |
| Action and outcome | Correspondence among instruction, action, and success label | Mark failure, recovery, and abort segments separately |
| Provenance and reproducibility | Generator, seed, source, transformation version | Link the history from data through to test results |
Eight diagnostic lenses at the level of the task episode. The fourth row, in bold, sits exactly where section 4 landed: you cannot decide what to augment until you know which condition is empty. These eight are a proposal for Physical AI practice, not a transcription of clauses from ISO/IEC 5259-2.
6.2Turn the generation goal into explicit conditions
Suppose the diagnosis concludes that failures cluster when the robot grasps a partially occluded cup. The next step is converting that sentence into generation conditions: degree of occlusion, the cup's material and position, approach direction, lighting, target location. Then measure again whether the generated data actually satisfies those conditions. One prohibition attaches here.
Do not treat a missing condition as filled merely because the new samples sit far away in embedding space. Distance in representation space is a proxy for diversity and says nothing about whether the condition you set out to fill was actually filled. The table in section 5 already put it plainly: visually clean data need not contain the right task-relevant variation for learning.
The generation history has to be recorded alongside. Nine items is the minimum: source episode identifier, generation method and model version, seed, environment conditions, coordinates and units, the provenance of the actions, how labels were produced, screening results, and the train/eval split. Without those nine, when the score later goes up there is no way to trace what raised it. The problem from section 5, where the model's contribution stays entangled with the optimizer, the reward model, and the data pipeline, repeats itself on the data side.
6.3From diagnosis to evaluation, one full turn
Join the pieces above and they form a single operating loop of six steps: gather observation, action, and outcome from field data; diagnose quality and locate the conditions where failures cluster; fix the conditions to augment and the screening criteria; confirm that the generated data actually filled those conditions; train policies under matched conditions and compare; measure again on the real robot. Without that last arrow the loop is not a loop. If failure conditions found in real evaluation do not return to the next diagnosis, the procedure is a checklist you use once and throw away.
The data quality operating loop. The two steps marked in orange are the ones most often skipped: augmentation without stated conditions, and evaluation whose results never return to the next diagnosis.
6.4Four arms, and the one that goes missing
Deciding whether synthetic data earned its keep takes comparison arms. Four are enough, and fewer than four leaves the question unanswerable.
| Arm | Training data composition | Question it answers |
|---|---|---|
| A. Baseline | Existing real data | What is current performance, and where does it fail? |
| B. Untargeted augmentation | A + synthetic data with no chosen target | Does adding synthetic data help at all? |
| C. Targeted augmentation | A + synthetic data chosen by diagnosis | At the same added volume, does choosing conditions pay? |
| D. Real augmentation | A + additional real data | How does it compare with real collection at similar cost? |
Four arms for judging the effect of synthetic data. The comparison has to hold training compute matched: record the number of update steps, frames or tokens processed, and GPU hours. More data changes the compute even at the same epoch count, so matching epochs alone does not make conditions equivalent.
Arm D is the most important seam in this report. It is the control arm missing from the experiment in section 4; it is the "matched training budgets and learners" condition the evidence table in section 5 demands of any synthetic data claim; and it is where practitioners ask how this compares with real collection at similar cost. Three separate sources point at the same single cell. When a synthetic data proposal lands on your desk, check first whether arm D is in the design.
6.5The model side needs its own questions
The eight items above diagnose data. Whether a model carries the evidence its claims require is a different question, and mixing the two blurs both. The position paper from section 5 proposes a minimum reporting set of three items for work involving physical robots. Recast as a buyer's questions:
- Did you measure action branching from matched resets? Start from the same state, feed in several different actions, and check whether action and outcome line up on low-dimensional quantities such as end-effector or object pose.
- Do three or four policies of clearly different strength come out in the right order? Look for calibration and ranking of closed-loop success rates, reported with confidence intervals.
- Did you actually execute the trajectory the model scored highest? This is a qualitative exploitability probe at minimum. Look for a record of checking whether the trajectory the model called best also succeeds on the real robot.
The third question is the empty cell from section 5. Since essentially nobody measures it, the answer will most likely be that it was not measured. Ask anyway. Once a question has a name, the next document gets a field for it.
One line on standards to close. The international standard specifying data quality measures (ISO/IEC 5259-2:2024) was published in November 2024, and on the manufacturing digital twin side, the US standards body frames verification, validation, and uncertainty quantification as a process that continues across the lifecycle rather than a one-time task. The eight items in this report are not a translation of those clauses, and they cannot be mapped one-to-one onto them. What can be said is that the direction is the same: quality is not measured once and filed away, it has to live inside the loop.
Why this matters to Pebblous
Everything so far converges on one shape. The average goes up while only certain cells move, and finding out which cells moved requires splitting the measurement by condition. That is why Pebblous has been watching this subject for a long time. Splitting a measurement by condition is what data diagnosis is.
7.1The same order: diagnose, augment, remeasure
Data Clinic's diagnosis and improvement proposals, its synthetic data generation, and PebbloScope's data exploration already sit in the same order as the operating loop in section 6: find the thin spots by diagnosis, augment against those spots, measure again under the same conditions. The reason Data Clinic's improvement cases name filling an underrepresented region and pruning duplicate or near-identical stretches as two different jobs has the same structure as the conclusion in section 4. What you add decides which cell rises, so adding and pruning have to come out of different diagnoses.
The boundary has to be stated plainly, though. Having that order does not mean current products automatically verify every sensor, action, and contact quality on a robot. Carrying a procedure that works on tabular and image data over to episode-level robot data still leaves open items, and among the eight lenses in section 6, physical validity and action-outcome correspondence are especially so.
7.2Data composition does not land just anywhere
Put what section 4 showed into data quality language and it comes out like this: the composition of training data predictably determines which capability cell it lands in. Placement-randomized trajectories landed in the placement condition; environment-varied trajectories landed in the distractor and unseen-object conditions. That means the chain running from data quality through the model's internal representation to real performance can be cut apart and measured condition by condition, which is the ground a diagnostic product stands on.
One open problem lives here too. When data is compared in representation space, whether that representation actually preserves the differences the task depends on has to be verified separately. Two scenes that look nearly identical can differ in mass or friction, and conversely two different backgrounds can demand the same motion. The prohibition from section 6, that distance in embedding space does not certify a filled condition, applies to the product side unchanged. Using representation-learning models for that verification is still a proposal under review inside Pebblous, not a feature shipped in an existing product.
7.3The three questions customers actually ask
Move all of this into a customer conversation and three questions remain. Which conditions in our setting can you reproduce? How can we trust the actions and labels in generated data? And how does it compare with real collection at the same cost? The third is arm D from section 6, the control arm missing in section 4, and the matched budget the evidence table in section 5 demanded. The cell three separate sources pointed at is the cell customers are already asking about.
A proposal is therefore better spent on the cost of failure and the verifiable scope than on a list of technology names. Which conditions can be reproduced today and which cannot yet; how far screening is automatic and where a human takes over. It helps to mark each item's status as available, planned for trial, or verified. A proposal that leaves those statuses unmarked ends up in the same position as the one-line headline from section 5.
7.4Where this work stands
Large generative models and physics engines can be brought in through open models or partner technology. Pebblous's place is therefore not in building a bigger generator but in condition-level quality diagnosis, choosing what to generate, screening output from multiple generators, comparing real effects, and managing the record of all of it. One assumption is worth refusing here: that model providers have no quality management of their own. They do, and they are expanding data processing and evaluation capabilities quickly.
The signal that demand is real sits in how the Korean program from section 4 divides its work. A program of roughly $23.4M (KRW 34B) assigned dedicated institutes to building the data collection platform and to data standardization with model validation. Diagnosis and validation were treated as roles in their own right rather than as add-on features. This is not a claim that Pebblous is part of that program; it is an observation that demand on the evaluation and standards side is already written into how the program was designed.
Sections 1 through 6 carry what could be verified in published papers, model cards, and announcements. This section 7 is the part those documents do not say. Please read them separately. The verbatim quotations and figures in the body were checked directly against full arXiv preprints, Hugging Face model cards, and official blogs, and the increments in the condition table and the metric breakdown were computed by this report from the original tables. Where a figure's underlying source could not be verified, the body says so on the spot. Thank you for reading a long article.
References
Sources differ in grade, so they are grouped. The first group is the material whose full text was obtained and checked line by line against the body and tables here. For the second group, bibliographic details and conclusions were cross-checked through the references of source 1, and only what that check covers is carried into the body. Vendor documents in the third group are cited with their self-reported character noted. The policy material in the fourth group could not be verified against an original press release, so it was cross-checked against multiple news reports, and the standards documents were checked only through their public summaries and scope, since the paid full texts were not opened.
The backbone of this report (primary full text checked)
- 1.Yang Yu, Shiyuan Zhang, Yifei Sheng, Haoxiang Ren, Haoxin Lin. "How Should World Models Be Evaluated for Embodied Decision-Making? A Decision-Making-Centric Position." arXiv:2606.15032 v2, June 28, 2026. Nanjing University. arXiv: 2606.15032 — CC BY 4.0. The six referents come from the abstract, the eight-level ladder and its questions from Table 7, the evidence standard for synthetic data claims from Table 6, the three census patterns with the authors' caveat and placement rule from §3.5, and the minimum reporting set for physical robots from §5.
- 2.Chenghao Gu and nine others. "GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions." arXiv:2608.06332 v1, August 6, 2026. arXiv: 2608.06332 — The experimental design and the verbatim policy-control statement in section 4 come from the body; the twenty condition-level values come from Appendix Table III. The increment table was computed by this report by plain subtraction across those three columns. There is no real-data control arm of equal size, and no sentence acknowledging that confound was found in the full text.
- 3.GigaAI, Tsinghua University. "GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation." arXiv:2607.02642 v1, July 2, 2026. arXiv: 2607.02642 — Apache-2.0. The six metric definitions and the 11.6% and 14.9% figures come from §6.5.1, per-metric values from Table 9, and the metric-group correlations and the static-video quotation from the findings in that same section. Equal weighting and contribution shares were derived after verifying that the plain mean of the six columns matches the printed average to four decimal places. The widely cited "RTX 4090" inference setup could not be found in the paper and is not used, and rollout-level grades are a different metric from the six-metric mean, so the two are not mixed.
- 4.Soroush Nasiriany and others. "RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots." arXiv:2406.02523, RSS 2024. arXiv: 2406.02523 — The quotation in section 3 is from the Limitations section, and it is the oldest evidence in which pipeline authors themselves record the need to screen generated trajectories.
Cross-checked academic literature
The works below were cross-checked for bibliographic details and conclusions through the reference list of source 1. What is carried into the body is limited to what each study reported; the full texts were not opened.
- 5.Yu Shang and others. "WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models." arXiv:2602.08971 (2026). arXiv: 2602.08971 — Section 5 cites its conclusion that perceptual rankings and functional rankings disagree.
- 6.Fanqi Jiang and others. "RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation." arXiv:2604.19092 (2026). arXiv: 2604.19092 — Section 5 cites its conclusion that video quality is not a reliable predictor of executability.
- 7.Yuhang Deng and eight others. "Rethinking Video Generation Model for the Embodied World" (R-Bench). arXiv:2601.15282 (2026). arXiv: 2601.15282 — The R-Bench entry in the section 5 census. Its authorship was reported two different ways; the reference list of source 1 independently confirms the attribution, and that same paper places this benchmark at L0 through L3.
- 8.Julian Quevedo, Percy Liang, Sherry Yang and others. "WorldGym." arXiv:2506.00613 v3, accepted at ICLR 2026. arXiv: 2506.00613 — The two sentences closing section 5 come from here: that policy rankings are preserved, and that generating realistic object interaction remains hard.
- 9.Meta AI. "V-JEPA 2.1." arXiv:2603.14482 v3, June 11, 2026. arXiv: 2603.14482 — The case closing section 2. The abstract's 20 points is the value for grasping, one of three tasks, and at a one-step horizon it is 10 points. The diagnosis that the remaining failures come from gripper action planning is the authors' own.
- 10.Danijar Hafner, Wilson Yan, Timothy Lillicrap. "Training Agents Inside of Scalable World Models" (Dreamer 4). arXiv:2509.24527, September 2025. arXiv: 2509.24527 — An entry in the section 2 landscape table, and the only one a year old, which is why its publication date is listed. The absence of official weight or code releases from the authors was checked against public repositories.
- 11.Saman Motamed and others. "Physics-IQ." arXiv:2501.09038 (2025) / Pranav Atreya, Karl Pertsch, Tony Lee and others. "RoboArena." arXiv:2506.18123, CoRL 2025 — Two entries in the section 5 census. The latter is the only one of the nine that compares policies on real robot hardware.
Primary vendor documents (model cards, official blogs, announcements)
- 12.NVIDIA. Cosmos 3 announcement (May 31, 2026, GTC Taipei) and the Cosmos3-Edge model card (Hugging Face
nvidia/Cosmos3-Edge, first posted July 20, 2026, updated August 25) — The spec and action-format tables in section 2 and the limitations quotation come from here. The announcement's "ranks first across…" and "months to days" are vendor self-reporting with no absolute numbers, and are not carried as fact. Only the model card's 4B parameter value is used; the technical report (June 22, 2026) could not be verified. - 13.World Labs. "Atlas" (September 1, 2026), "Building Worlds That Train Robots" (July 28, 2026), and the Marble and World API documentation — The release status and the robotics quotation in section 2 come from here. Access is partner-only early access, with no weights, public API, pricing, or performance figures. The statement that no limitations passage could be found is the result of checking these documents. Marble pricing is not used because secondary sources disagree.
- 14.Google DeepMind. Genie 3 research announcement (August 5, 2025) / Google Labs. Project Genie (released January 29, 2026, expanded in May) — That it is not a game engine, and the 60-second, 24fps, 720p limits, are the published constraints. Subscription pricing is not used because secondary sources disagree.
Policy and standards
- 15.Ministry of Science and ICT and the Institute of Information and Communications Technology Planning and Evaluation (Korea). Kickoff meeting for the Physical AI leading-technology development program (June 9, 2026, LG Science Park, Magok, Seoul) — The budget, duration, lead organization, participants, number of validation rounds, and the core-goal quotation in section 4 come from here. The original press release could not be verified, so the details were cross-checked against multiple news reports from June 9 and 10, 2026, and the 14.5 percentage points cited as a comparison is carried only as something the program materials cite, since its underlying source could not be found. The principal investigator's remarks are likewise a secondary confirmation.
- 16.ISO/IEC 5259-2:2024, "Artificial intelligence — Data quality for analytics and machine learning (ML) — Part 2: Data quality measures" (JTC 1/SC 42, November 2024) / NIST, "Digital Twins for Advanced Manufacturing," on verification, validation, and uncertainty quantification — The closing paragraph of section 6. The paid standard's full text was not opened, so only the public summary and scope were checked, and no clause is quoted.
Related pieces on the Pebblous blog
- 17.World model survey (section 1), Physical AI industry landscape (section 2) — The first covers the academic taxonomy, the second the industrial axis.
- 18.Synthetic data pipeline, Physical verification of world model output (section 3), Robot learning data layers (section 6) — The three threads this report links out to rather than repeating.