Appearance
Grounded Embodied AI: Contact, Constraints, and Collision-Testing
The same wall, three robot arms
Embodied AI keeps failing in the same place: at the boundary between abstraction and physics. A VLA model can narrate how to rotate a valve cap, then fumble the rotation because the gripper occludes the valve and the model's world model lives in pixels, not forces. A construction robot following a predetermined block sequence loses the arch when a 3D-printed block is 2 mm off spec. An AV planner sails a clean simulated route while a real test pipeline never throws the blind-spot pull-out scenario at it, because the steps needed to generate, modify, execute, and analyze that scenario live in four unrelated tools.
Three papers crossed my desk this month, each attacking that boundary from a different domain. DeCAL adds tactile grounding to dexterous VLA models. HSAC deletes the construction blueprint and replaces it with stability-aware reinforcement learning. PlannerForge collapses the entire scenario-based testing pipeline for autonomous driving into one LLM-agent framework. Different robots, same thesis: rigid abstraction breaks, physical feedback doesn't.
DeCAL: when pixels aren't enough
Dexterous manipulation is where VLA models struggle most. The robot's hand covers the target object, so the camera sees occlusion right at the moment when grasp details matter. Contact happens at unpredictable points, so the dynamics that mattered in the last trial don't transfer to the next one.
DeCAL is a physically-grounded dexterous VLA built on a Mixture-of-Transformers (MoT) backbone. The idea is to split the model into three jobs: understanding, imagination, and action generation, run by specialized experts, then let information flow between them instead of forcing one monolithic network to do all three. Each expert keeps its own capacity, and MoT routing decides how their representations interact.
Two components carry the weight. Adaptive Visuo-Tactile Fusion uses a contact-aware gating strategy to decide when tactile information gets to influence the network, rather than blending vision and touch at every timestep with fixed weights. Visuo-Tactile Latent Co-Imagination trains the model to predict visual and tactile dynamics jointly in a shared latent space. That second part is where the physical grounding happens: the model doesn't just imagine where the object will be, it imagines what the next contact will feel like.
The results support the design. DeCAL reports a 71% average success rate across its task suite, the best published results on these dexterous manipulation benchmarks, and an 83.4% progress success rate, which measures how far the policy advances a task it hasn't seen in training.
What actually changed: gating and physical imagination
The two ideas behind DeCAL look simple on paper, but they answer a specific set of failures that plague contact-rich VLA work:
| Failure mode | Why it happens | DeCAL's counter |
|---|---|---|
| Severe occlusion | The hand and target block the camera at the worst moment | Tactile readings need no line of sight, and gating lets them dominate when contact starts |
| Contact dynamics ignored | Trajectory prediction in pixel space has no force model | Latent co-imagination predicts visual and tactile dynamics jointly, giving the policy an implicit physics model |
| Tactile signal diluted | Touch is only informative at contact instants, not in between | Contact-aware gating suppresses touch when it's noise and amplifies it when it's signal |
The gating point is the subtle one. Naive multimodal fusion treats every tactile frame as equally informative, which means the rare moment of actual contact gets averaged into a pile of idle sensor noise. DeCAL's gating is a learned decision about when touch deserves attention. That's the difference between bolting a sensor on and making the policy reason about the sensor's informativeness.
Quick Take: All three papers replace some measured abstraction with a direct physical signal, contact force, structural stability, or collision outcome.
Building without blueprints
Robotic construction raises the opposite problem. Nothing is occluded and the geometry is simple, yet current methods fail because they plan rigidly. High-precision plans can't absorb the tolerances, inaccuracies, and unexpected changes inherent to physical fabrication.
The new paper introduces HSAC, a reinforcement learning approach that abandons predefined plans entirely. The policy generates the construction sequence adaptively as the structure is built, which requires two things a standard RL setup doesn't have. First, a graph-structured state representation, where blocks are nodes and edges encode support relationships. Second, a mixed action space, because choosing a block is discrete but placing it is continuous.
The technical twist is the "unilateral edges" in the graph neural network. A unilateral edge encodes that block A resting on block B is stable while the reverse is not, a directional support relation. That makes the stability signal cheaper to compute, which matters because simulating the stability of a partial structure is the computational bottleneck. HSAC extends soft actor-critic (SAC) to this hybrid discrete-continuous setting. The authors report comfortably higher asymptotic performance than the prior hybrid-PPO (HPPO) method, with good sample efficiency and robustness to hyperparameter choices, which is rarer than it should be in hybrid-action RL.
Key numbers for HSAC
- Handles up to 10 discrete actions per step without performance degradation, roughly the choice set of a small block palette
- Transfers sim-to-real: a physical two-robot setup built a spanning arch from 3D-printed blocks in closed-loop execution
- Zero predefined plans: block selection and placement come from the stability graph, not a script
The spanning arch demo is the result that matters. A policy trained entirely in simulation, with no fine-tuning on hardware, picked and placed real 3D-printed blocks until the structure stood. Closed-loop execution means the policy saw each placement result and adjusted the next one, absorbing the tolerance errors that would have killed a plan-based counterpart.
The testing pipeline gets one brain
Autonomous driving has a parallel rigidity problem in its tooling. Scenario-based testing is the standard way to validate ADS behavior, but it's a fragmented modular pipeline: scenario generation, retrieval, modification, ADS execution, and results analysis each run in separate tools with barely any interaction. An engineer who wants to ask "what happens if a car cuts in from a blind spot at a tunnel exit?" has to hand-craft the scenario, translate it into the simulator's format, run the ADS, then manually inspect logs.
PlannerForge is an LLM-agent framework that covers the whole pipeline, from Scenario Generation to ADS Assessment, and adds two stages a modular pipeline can't easily have: ADS Enhancement and ADS Benchmarking. Since every stage is an agent, information flows backward as well as forward. The planner-testing stage can tell the scenario-generation stage which scenario shapes actually stress the planner, and the enhancement stage can modify the planner configuration and re-run the same scenarios.
The evaluation is thorough. The authors ran 10 off-the-shelf LLMs across all tasks under 5 prompt conditions. Best-per-task scores range from 0.88 to 1.00, and the striking result is that open-source 20-35B models match commercial APIs on most tasks. Qwen3.6:35B, roughly the size that fits on one A100-class GPU at 16-bit, matches commercial backends on three of the five tasks.
That generation gap translates directly: 49 more executable scenarios out of 200 is the difference between a test suite with holes in it and one you can ship. The cost-tuning result is the most practically useful finding. At N=400 scenario executions, PlannerForge lifted planner success from 50.4% to 70.2% and cut collisions from 19.0% to 8.4%, without any domain-specific fine-tuning. The LLM agents found weak spots in the planner, suggested configuration changes, and re-tested. 400 simulated executions is cheap next to a single road-test day, which is what makes loop-closed testing attractive to teams that can't burn track miles.
There is a cost to chaining LLM agents: the pipeline retains 83% of seed queries with commercial models and 78% with open models. Agentic orchestration leaks queries at module boundaries. Those numbers are still high enough to be useful, but they're a reminder that a chain of LLM calls is lossy.
The convergence: feedback beats fidelity
Strip away the domains and all three systems make the same move. Each one takes a signal the physical world actually provides, contact, stability, collision, and closes a control or reasoning loop around it.
| DeCAL | HSAC | PlannerForge | |
|---|---|---|---|
| Domain | Dexterous manipulation | Robotic construction | AV motion-planning validation |
| Physical signal | Tactile contact | Structural stability | Executable crash scenarios |
| Core architecture | VLA on a Mixture-of-Transformers | Graph RL, SAC extension | LLM-agent pipeline chaining tools |
| What it removes | Pixels-only trajectory imagination | Predefined construction plans | Fragmented modular toolchain |
| Headline result | 71% average success, 83.4% progress | Two-robot spanning arch on real hardware | 193/200 executable scenarios; planner success 50.4% to 70.2% |
The older approach in each domain tried to make the abstraction better: better simulation fidelity, better plan search, better scenario templates. These papers all chose a different path. They made the model listen to physics at the moment it matters, and they let the model adjust as physics responds. That's the pattern worth copying, regardless of your domain.
Common Pitfalls (what trips people up)
The papers are explicit about what goes wrong, so here are the traps in concrete form.
Don't fuse tactile data homogeneously. If you concatenate touch readings with vision tokens at every timestep, the rare informative contact is drowned in idle sensor noise. Gate the tactile stream on contact events, either with a learned attention mechanism or a heuristic contact detector, and let touch dominate only when it's informative.
Don't trust sim-trained construction policies to replay sequences. HSAC's arch build worked because the policy closed the loop on the observed graph state after every placement. A plan-based policy trained in simulation would accumulate tolerance error with each block, and 3D-printed parts guarantee tolerance error.
Don't conflate a plausible text scenario with a valid crash scenario. From-Words-to-Collisions, the prior text-to-scenario system, produces physically valid edits just 31% of the time, while PlannerForge clears 94%. If you're generating test scenarios with an LLM, run every output through the physics simulator before you trust it. Text sounding realistic is not the same as a car existing at that position.
Don't assume the biggest LLM is the best test agent. The 20-35B open-weights class matched commercial APIs on most testing tasks, and prompt conditions moved scores as much as model choice did. Price out a Qwen3.6:35B deployment before buying API credits for a larger commercial model.
Don't design the testing pipeline as modules and hope chaining works for free. End-to-end retention of 78% with open models means roughly one in six seed queries vanishes between pipeline stages. If you don't engineer the module interfaces to pass queries through, you're testing a filtered subset of your intended scenario space, and you won't know which subset.
One thing to remember
Physical grounding is the through-line. DeCAL's contact-aware gating, HSAC's unilateral stability edges, and PlannerForge's executable-scenario validation all do the same job: they force the abstraction to pay rent to reality. When you design an embodied AI system, the question to ask isn't "how do I make my model smarter?" but "what physical signal is my model currently ignoring, and how do I make it impossible to ignore?"
The Bottom Line
If you're building dexterous manipulation policies, add tactile sensing with contact-aware gating and joint latent-dynamics modeling, because vision-only imagination fails exactly when the hand covers the target. DeCAL's 83.4% progress rate on unseen scenarios is the payoff for that grounding.
If you're deploying robots for assembly or construction, drop the predefined plan and train a stability-aware policy on graph state in simulation. HSAC's transfer to a real two-robot arch build, with zero hardware fine-tuning, shows that closed-loop policies absorb tolerance errors that open-loop plans can't.
If you validate AV planners, consolidate scenario-based testing into a single LLM-agent framework and start with a 20-35B open-weight model, since it matches commercial APIs on most tasks and keeps 78% of seed scenarios through the chain. One thing to watch: LLM-agent testing pipelines are moving from research to tooling fast, and they'll likely become the default ADS validation setup within a year.