Skip to content

What Actually Breaks When Learned Policies Hit Real Robots

#robot-learning #imitation-learning #sim-to-real #contact-rich-manipulation #multimodal-llm #drone-control #reinforcement-learning

The Robustness That Evaluation Misses ​

A policy that hits 100% task success in the lab can collapse on the factory floor, at a faster clock speed, or the moment it has to declare "I'm done." Four recent papers each isolate one of those hidden failure modes. Read together, they map where learned robot policies actually break.

The pattern: evaluation measures averages, deployment punishes specific failures. Facet-0 attacks contact outcomes you can't see from task progress alone. ParcelStow shows temporal robustness is its own axis, separate from scene or object robustness. DroneCATS finds navigation is easy and knowing when to stop is hard. The safe sim-to-real work adds that even the data you collect to fix these gaps carries a safety price.

82% mean success on five sub-millimeter assembly tasks, vs 15% for the strongest baseline 0.5 mm placement accuracy at 50 ms command latency (Facet-0) 100% → 53% ACT success from nominal to max demonstrated speed, vs 100% → 84% for the scripted expert 0 of 414 acquisitions without force closure completed the task 2B parameters is enough to navigate a drone correctly, but not to declare arrival properly

Contact Consequences as a First-Class Signal ​

Sub-millimeter assembly is where vision-only policies stop working. When a part seats at 0.5 mm tolerance, the useful information lives in the wrist force and torque, the wrench, not just the pixels. Two motions can make identical task progress and produce completely different contact outcomes: one is seating the part, the other is jamming it. A policy that only predicts the next pose has no way to tell them apart.

Facet-0 (arXiv:2609.01596) treats the wrench as something to predict and value. The architecture aligns a causal wrench history with vision-language semantics and kinematic state, then uses flow matching to generate each action chunk together with the future wrist-wrench profile that chunk should induce. A distributional Action-Wrench Critic scores motions by contact outcome rather than task progress. Phase-aware rewards and contact-selective credit push the policy to concentrate on decisive interactions.

The design gets practical in the adaptation layer. A lightweight bounded actor sits on top of a frozen representation, so on-robot adaptation for part-specific dynamics doesn't destabilize the learned contact model. RL stays defined over executable Cartesian actions, while an auxiliary wrench head keeps predicting non-commanded contact coupling.

The numbers back the approach. 82% mean success on five computer-assembly tasks versus 15% for the strongest baseline, at 0.5 mm placement accuracy. Command latency of 50 ms is fast enough to react mid-insertion instead of after it.

Temporal Robustness: Imitation Copies the Average, Not the Range ​

Imitation evaluation usually varies the scene, the object, the instruction. It rarely varies the clock. ParcelStow asks a narrower question: under identical conditions, does the learner keep the expert's robustness across execution speeds?

Here's the setup. A scripted expert and an ACT policy trained on its demonstrations both hit 100% task success at nominal speed. ParcelStow is contact-rich: acquire, reorient, insert a parcel. Demonstrations cover the speedup range for the phases after acquisition. At the maximum demonstrated speed, the expert still succeeds 84% of the time. ACT drops to 53%. Two ACT policies with different initializations degraded by 34 and 48 percentage points. The expert lost 16.

Insertion is where it breaks. 35 of ACT's 47 failures at max speed were insertion misalignments. And the force-closure result is brutal: across all policies and speeds, none of the 414 acquisitions without force closure completed the task.

ACT isn't a bad policy. The real finding is that nominal success is almost uninformative about temporal robustness. Two policies can look identical in the eval suite and diverge by 31 points the moment the clock changes.

Quick Take: If your evaluation only runs at nominal speed, "100% task success" tells you nothing about whether the policy survives a faster production cycle.

The Action Protocol Is the Hard Part ​

DroneCATS does something stubborn: it drops an MLLM directly into the drone's control loop, declares the entire action space in the prompt, and refuses to narrow the model's job. The model has to yaw, search, deliberate when unsure, and self-declare arrival. No fine-tuning, no function-calling schemas. The benchmark treats the model as the independent variable and scales the roster down to 2B parameters.

The paradox: flying isn't what fails. Small open models navigate into the success radius more reliably than frontier models, then lose the episode by declaring arrival prematurely or not at all. Multi-drone commanding amplifies the divide: small models crib a single coordinate across distinct views instead of issuing per-fleet commands.

That's a sharp result. The model's spatial perception holds up, but its action protocol doesn't. For edge deployment this matters more than FLOPs: a 2B model that plans persistently and knows when it's done beats a frontier model that navigates well and terminates badly. The terminating action (arrival, task complete, "done") is a decision, not a side effect, and it can be the entire failure.

The Data You Need to Fix the Gap Has Its Own Price ​

The first three papers assume you can collect real-world data to diagnose and fix failures. The fourth asks what happens when that collection is itself dangerous. Direct sim-to-real transfer isn't guaranteed to work, but naive real-world RL to correct the sim gap can damage the system you're training. In robotics and healthcare this is the constraint, not an edge case.

The paper formalizes safe sim-to-real transfer as reward-free safe RL. The agent exploits the imperfect simulator to reduce real-world interaction, but real-world exploration must stay safe, and the output must be a near-optimal feasible policy for any reward function that might arrive later. The real-world sample complexity bound characterizes how much interaction you need as a function of sim-to-real mismatch. Large mismatch, and you pay close to pure real-world RL. Small mismatch, and the bound shrinks proportionally.

Practical read: the value of your simulator is a number you can estimate, and the safety budget belongs inside the algorithm design, not bolted on afterward.

Four Failure Modes, One Lesson ​

When I line the four papers up, the through-line is obvious.

PaperWhat breaksCore mechanismHeadline result
Facet-0 (2609.01596)Contact outcomes invisible to pose predictionAction-wrench critic, flow matching over wrench profiles82% vs 15% success, 0.5 mm placement
ParcelStow (2609.01453)Temporal robustness when speed changesExpert-learner eval across speedup factorsACT 53% vs expert 84% at max speed
DroneCATS (2609.01404)Termination and multi-agent protocolPrompt-declared action space, no fine-tuning2B models navigate but declare arrival badly
Safe sim-to-real (2609.01418)Safe real-world data collection under mismatchReward-free safe RL, provable sample boundsReal samples bounded by sim-to-real mismatch

Every one of these is a robustness axis that doesn't show up in the headline number. Task success at nominal conditions is the wrong single metric for embodied policies. The useful question is always the same: what breaks first at the edge?

What Trips People Up ​

  1. Treating force/torque as optional telemetry. If your insertion policy can't distinguish equal-progress motions with different contact outcomes, it's flying blind. That distinction is the whole reason Facet-0's critic exists. Feed the wrench in and predict it, don't just log it.

  2. Reading 100% at nominal speed as done. ParcelStow shows two policies that are identical at nominal speed and 31 points apart at max speed. Put execution-speed variation in your eval suite, and test faster than your demo collection speed.

  3. Treating the terminating action as a detail. DroneCATS shows a model navigating into the success radius and then losing the episode by declaring arrival early or not at all. If the agent can't reliably emit "done", the whole run fails at the last step. Evaluate termination explicitly.

  4. Advancing phases without a force-closure check. Zero of 414 acquisitions without force closure completed ParcelStow. If your pipeline moves to the next phase without verifying contact state, it's betting against the data.

  5. Collecting real-world correction data with no safety plan. The safe sim-to-real paper formalizes what practitioners feel: the worse the simulator, the more real interaction you need, and you need that budget known up front. Estimate the mismatch before you let the policy explore.

Sources ​

One Thing to Remember ​

The headline success rate hides where the policy breaks. Each of these papers found its failure only after adding one axis the benchmark was missing: contact consequences, execution speed, termination protocol, or data-collection safety. Design your evaluation to find the first failure, not to confirm the average.

The Bottom Line ​

If you're building contact-rich manipulation policies, treat the wrist wrench as an input and a predicted output. Equal task progress with different contact outcomes is the failure mode vision alone can't see.

If you're doing imitation learning for deployment, add execution-speed variation to your evaluation before you trust a 100% nominal success rate. An expert and learner that look identical at nominal speed can diverge by 30 points when the clock changes.

If you're putting an MLLM on an edge robot, optimize for protocol discipline over raw perception. A 2B model that reliably knows when it's done beats a frontier model that navigates perfectly and terminates badly.