Executive Summary
Three researchers at the School of Artificial Intelligence, Jilin University, posted a paper to arXiv on September 25, 2026. On top of LIBERO, a simulator for robot manipulation learning, they built 2,826 tasks whose instructions are themselves false, and put eight models that read a camera view together with a human sentence and move a robot arm through them one by one. Pick up a cup that is not on the table, open a drawer that is already open, clear away the plate that is buried underneath first. Those are the instructions. This article looks at what a robot given that kind of instruction actually did.
Every success rate fell, and the biggest of those losses was OpenVLA's 56.2 percentage points. The paper puts its weight somewhere other than that number. When the attempts recorded as failures are opened again, the robot turns out not to have stopped. It kept extending the arm toward the original goal, its first twenty steps were nearly indistinguishable from a normal run, and it barely reduced the force of its motions. The authors gave this behavior a name: Failed Persistence.
Sections 1 through 4 report what is in the paper. Section 5 reads those same facts through the lens of data quality, and that reading belongs to this article.
Key Figures
Source: Hou, Wu, Chang, "ConflictVLA-Bench" (arXiv:2609.31792, September 25, 2026).
56.2 pts
Drop in OpenVLA's original-goal rate
78.7% on valid instructions, 22.5% on false ones
17.3 pts
Smallest drop among the eight
π₀.₅'s figure. No model held steady under false instructions
0.845
Early-path similarity of π₀-FAST's failed runs
Half of them ended in failure, yet the first twenty steps looked almost normal
0 models
Moved as intended on all four metrics by the premise-check prompt
Not one of the eight
What a Robot Does With an Impossible Instruction
A VLA is a model that reads the scene from a camera together with the sentence a person hands over, and puts out the robot arm's next motion directly. After Stanford researchers released OpenVLA as an open model in 2024, names such as the π₀ family, GR00T and UniVLA followed one after another. The places where these models have had their skill measured all stood on a single assumption: that the instruction a person gives is correct.
A real room is not like that. Someone can ask for a cup that is not on the table, ask for a drawer that is already open, or ask for the plate underneath to be cleared away first. What is wrong is not the side receiving the instruction but the instruction itself. The paper calls this a premise conflict and splits it four ways.
- Instruction-internal conflict — the goal and the constraint inside one sentence contradict each other
- Object-grounding conflict — the object the instruction points at is absent from the scene or cannot be pinned down
- Spatial-relation conflict — the positional relation the instruction takes to be true is false in the scene
- Physical-feasibility conflict — the instruction demands a motion or an order that cannot physically be carried out
There is one more axis. The same error planted in a task that ends with a single grasp and planted in a task that has to be stepped through are different experiences for the robot, and an error concentrated in one spot is different from one scattered about. The researchers multiplied these two splits and sorted the tasks into four configurations, L1 through L4. One error in one step is L1; several errors in one step touching the same requirement together is L2; a single sub-goal carrying the error among several steps is L3; entangled errors spread across several steps is L4. The paper nails down in advance that this is not to be read as a difficulty order. These configurations describe construction structure rather than an assumed empirical difficulty ordering, the authors write, and the result tables bear that out by reporting no breakdown across the four.
The floor these tasks sit on is LIBERO, a robot manipulation benchmark released in 2023 where every task carries a machine-readable record of the objects placed in the scene and the conditions for success. The researchers read those specifications to pull out objects and predicates, planted conflicts in them, paired each one with the original valid task, and then ran automatic checks followed by human review. The review had two stages. Tasks where the conflict did not hold or could not be confirmed were thrown out, and tasks whose notation was wrong but fixable were fixed and checked again. What survived was 1,413 base tasks. With two prompt conditions applied, the count becomes 2,826.
The pairing is the heart of this design. Every rollout run under a false instruction has one valid rollout attached to it, in the same scene and with the same goal. With a baseline sitting right beside it, the question is no longer only whether the task succeeded but whether this robot moved differently from usual. A single outcome cell cannot answer that.
Here is the scale of the experiment. Each of the 2,826 tasks was run from five different initial states, and the 40 valid tasks that serve as the reference were run from the same five states under both prompt conditions. That comes to 14,130 conflict rollouts and 400 valid rollouts per model, 14,530 in all. Every percentage below counts those rollouts. The figure of 40 reference tasks is hard to pass over. The 1,413 conflict tasks are not 1,413 different scenes; they were grown on top of LIBERO's 40 base tasks by varying the error planted in them.
Taken from the outcome side, OpenVLA is the one that fell furthest of the eight. A success rate of 78.7% on valid instructions came down to 22.5% under false ones. π₀.₅, which fell least, still lost 17.3 percentage points, from 94.5% to 77.2%. A strong ordinary score was no protection either. VLA-JEPA put up 99.0% on valid instructions and settled at 69.1% under false ones.
The raw figures per model are below. The left column is the success rate on valid instructions, and the middle one is the share of runs under a false instruction that still completed the original goal.
| Model | Valid-instruction success | Original goal under conflict | Drop |
|---|---|---|---|
| OpenVLA | 78.7% | 22.5% | 56.2 pts |
| π₀-FAST | 83.4% | 44.8% | 38.6 pts |
| π₀ | 93.1% | 57.6% | 35.5 pts |
| VLA-JEPA | 99.0% | 69.1% | 29.9 pts |
| GR00T N1.7 | 92.1% | 63.2% | 29.0 pts |
| OpenVLA-OFT | 98.3% | 71.1% | 27.1 pts |
| UniVLA | 84.1% | 59.5% | 24.6 pts |
| π₀.₅ | 94.5% | 77.2% | 17.3 pts |
Failed Persistence: The Behavior Hidden by a Success Table
That much is the success-rate story, and what the paper really wanted to ask comes after it. The single cell marked failure records two entirely different events in the same shape. In one, the robot noticed that the instruction was odd and folded up its motion. In the other, it never noticed and pushed toward the original goal until it fell over. On a screen both are failures, but on a machine with an arm and a gripper the two do not weigh the same.
The authors named the second behavior Failed Persistence. The definition is a single sentence: the reference goal is not completed, but observable reference-task behavior remains during execution.
To split the two branches, the researchers brought out three process metrics. All of them are relative values measured against the valid rollout attached alongside.
- Target Approach Ratio (TAR) — how far the gripper tip actually closed in on the original target object
- Trajectory Similarity (TS) — how much the path the arm drew overlaps the path of a normal run, measured with the time axis stretched and squeezed to match
- Action Magnitude Suppression (AMS) — how much the force of the motions was actually reduced, the reduction that should have come if the anomaly had registered
The eight-model averages lean one way. Target approach sits between 0.765 and 0.896, and path similarity over the first twenty steps alone between 0.784 and 0.886. Action suppression stopped at 0.067 to 0.276. That combination, two high values against one low one, says almost everything this experiment concluded. The robot went on toward the original goal much as it usually does, and it did not slow down. Over the whole trajectory, similarity drops to a range of 0.471 to 0.744, which means the departure from the normal path arrived late.
Early similarity running higher than whole-trajectory similarity held without exception across the eight models. The start resembled a normal run, and the divergence piled up afterward. The authors take no further step here. They write that the measurements cannot identify the internal cause of this divergence, but that they are consistent with the policy first initiating familiar task behavior and later accumulating differences during localization, grasping, or control, and there they stop.
Pulling out only the failed attempts sharpens the picture. The paper calls a pair capability-qualified when the valid side succeeded, and gathered again only those that failed under a false instruction. That clears away in advance the business of counting a task the model could never do as a failure. Even in this subset, target approach held at 0.586 to 0.741 and early path similarity at 0.653 to 0.845, while action suppression stayed at 0.133 to 0.370. A good share of the records marked failure were in fact runs that had never let go of the original goal.
The model the paper holds up as an example is π₀-FAST. It lost 50.2% of its capability-qualified pairs to failure. By the score alone it is a model that fails about half the time. Yet the target approach of those failed records is 0.741 and their early path similarity 0.845, while action suppression stopped at 0.178. This is not a model that fails half the time. It is a model that almost always pushes to the end and then fails, and that character would have stayed invisible in the success-rate column.
Nothing in the paper puts this point more briefly than the closing line of its abstract. Terminal failure alone, the authors write, establishes neither behavioral disengagement nor refusal, and outcomes alone are insufficient for VLA evaluation.
Telling the Model to Check the Premise First Does Little
This is where the count of 2,826 tasks being twice 1,413 comes from. The researchers ran the same tasks under two prompts. One gave the plain instruction only. The other laid a sentence in front of it asking the model to check first whether the premise holds. The design tests whether a remedy that often works on the language-model side also works on a model with a body.
The result comes down to four numbers. Target approach went down in four of the eight models, early path similarity in three, whole-path similarity in four, and action suppression went up in four. Not one model moved in the intended direction on all four metrics. In some, one column improved while another got worse. In others, nothing responded anywhere.
Worse things happened too. When OpenVLA and π₀-FAST received the sentence telling them to check the premise, even their success rate on valid instructions fell, by 20.2 percentage points for OpenVLA and 38.8 for π₀-FAST. They did not turn cautious only in front of false instructions. Their hands stiffened on perfectly good ones as well. The authors read this as general prompt-induced disruption rather than selective regulation under conflict.
The paper's own sentence reads this way. Explicit instructions to check task premises do not consistently produce selective and coordinated behavioral changes. It means today's VLA holds no premise-verification structure inside it, and that one line of text cannot build that structure in.
Robots Break Down Worst on Space and Physics
The four conflict types did not leave the same mark. The figures below are all averages over the eight models weighted equally, laid out metric beside matching metric so that the types can be compared.
Target approach first. An instruction with an internal contradiction gives 0.938, and a mistaken object reference 0.923. A false spatial relation comes down to 0.803, and a physically impossible demand to 0.795. In the first two cases the robot headed for the original goal even more squarely than elsewhere. A contradiction inside a sentence, or an object that is not there, barely disturbed the movement of the arm.
The outcome-side numbers point the other way. Under instruction-internal conflicts and object-grounding conflicts, the rate of completing the original goal stayed comparatively high at 69.9% and 64.3%. Under spatial-relation conflicts and physical-feasibility conflicts, the drop in that rate reached 38.7 and 45.0 percentage points. The more directly a conflict collides with the physics of the scene, the more both the outcome and the process shake.
Laid on top of each other, the two observations produce one worrying combination. A contradiction inside the instruction and a missing object are the easier kinds to catch, and yet that is where the robot marched toward the goal most calmly, while a physically impossible demand damaged the outcome most. Behavior does not stop at the errors that could be checked in words, and the loss runs largest at the errors that only show up when a body runs into them.
On safety the authors are plain. Because this kind of motion acts directly on the physical environment, continued execution can damage objects, cause unintended manipulation, or violate safety constraints.
Why Pebblous Is Watching This Research
From here we read the same numbers from the side of our own workplace.
The material behind this benchmark comes first. The 2,826 tasks are not newly filmed data. They were made by reading task specifications that already existed in LIBERO and planting errors in them. Training and evaluation data usually brings to mind gathering more of it, and cleaner. This paper chose the other side. There are properties that stay invisible unless someone builds pairs with a deliberately false premise planted in them. This is not the accidental contamination of missing values or noise. It is designed contamination.
Next comes what gets measured. The habit of speaking about quality through a single success rate is lodged deep in data work as well. A score sheet written as so many percent accurate and so many percent in error is easy to read and handy in a meeting. The trouble is that the one cell erases the cause. In this paper the cell marked failure put the robot that stopped and the robot that pushed on in the same place, and what separated them was the process-side record: target approach, path similarity, action suppression. Data quality is no different. Why this cell is empty, and at which stage this label flipped, come out of the traces left in the processing rather than out of an outcome score.
The prompt experiment left a third lesson. One line asking the model to check the premise fixed none of the eight, and two of them lost performance they had been holding. That puts a number on how the attempt to buy safety with an instruction ends. The result also shows, in reverse, where the work belongs. Situations with false premises have to go into training and evaluation data as a formal category, and what happened in those situations has to be recorded down to the process. Both fall to the side that designs the data.
Last, this paper drew a boundary around itself, and that deserves a note. The experiment took place inside a simulator, and nothing anywhere guarantees the same figures on a real robot. The count of 2,826 tasks grew along the variety of errors rather than the variety of scenes, so whether the same picture appears outside the 40 base tasks remains unknown to everyone. The authors did not widen their conclusion into an absence of refusal capability. They stopped at the line that terminal failure alone does not establish refusal. A sentence that claims only as much as it can claim tells the side quoting it what else needs checking.
Thank you for reading this far. The paper itself and the experiment materials are available on arXiv and at the project page. If you count how many of the metrics your team uses to talk about the quality of a model or a pipeline carry process rather than outcome, we would be glad to hear what turned out to be missing from that list.
References
R.1Academic Papers
- 1.Hou, L., Wu, Y., & Chang, Y. (2026). "ConflictVLA-Bench: Benchmarking Behavioral Responses of Vision-Language-Action Models to Premise Conflicts." arXiv:2609.31792. School of Artificial Intelligence, Jilin University. Unless noted otherwise, every figure and quotation in this article comes from here: the task construction procedure, the per-model success-rate table, the definitions and values of the process metrics, the comparison across prompt conditions, and the results by conflict type.
- 2.Liu, B., Zhu, Y., Gao, C., Feng, Y., Liu, Q., Zhu, Y., & Stone, P. (2023). "LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning." arXiv:2306.03310 (NeurIPS 2023 Datasets and Benchmarks). The robot manipulation learning benchmark whose task specifications and scenes the work above borrowed.
- 3.Kim, M. J., Pertsch, K., Karamcheti, S., et al. (2024). "OpenVLA: An Open-Source Vision-Language-Action Model." arXiv:2406.09246. The 7-billion-parameter open model this article's main search keyword points to, trained on 970,000 real-robot demonstrations.
R.2Released Materials
- 4.ConflictVLA-Bench project page. github.com/EmbodiedAISurvey/ConflictVLA-Bench. Where the paper posted its experiment data and supplementary materials.