Appearance
World models are data gluttons
World models are having a moment in embodied AI. The pitch is simple: learn how the world evolves from pixels and actions, then plan or train policies inside that learned simulation. NVIDIA's Cosmos, the π0 series, GR00T N, and a growing list of startups are all betting on it.
But there's a catch. World models are data gluttons, and the data they need doesn't exist at the scale they need it. Teleoperation is slow and expensive. Internet video is abundant but wrong in subtle ways. Every lab is solving this differently.
A cluster of recent releases shows the field attacking the problem from four directions at once. NVIDIA shipped a cleaned, standardized version of the DROID teleop dataset. LightwheelAI opened a gated 90,000-hour egocentric video dataset with hand-pose annotations. A new paper, Game2World, showed how to strip UI from gameplay footage so it's usable for training. And Veeda AI, founded by ex-NVIDIA researchers, raised $90M to generate training environments with world models themselves.
They're layers of the same stack.
Key numbers: 71,907 DROID episodes now in LeRobot v3.0 format (707 GB). 90,000 hours of egocentric video in EgoStandard. 96K synthetic gameplay clips in Game2World. An order-of-magnitude planning speedup from LeFlow. $90M seed for Veeda's generative simulation bet.
The data hierarchy driving everything
The 36kr piece on Veeda lays out the data hierarchy the whole field is working with. At the top: teleoperation, high quality but expensive enough that you can only collect hundreds of hours. Below that: sensor data from UMI-style grippers and wearable cameras, cheaper and more abundant but with a weaker action signal. Below that: generated data from world models and games, nearly unlimited but needing verification. At the bottom: raw internet video, free but noisy.
The result is a compromise everyone is stuck with: pretrain on cheap data, post-train on expensive data. It works, but it's a patch.
Each tier is changing fast, in very different ways.
Gameplay video: cleaning the cheapest tier
Gameplay footage is the most abundant source of interaction data on the internet. Millions of hours, diverse environments, complex dynamics. The problem: raw gameplay entangles the game world with screen-space UI. Health bars, minimaps, crosshairs, inventory slots. A world model trained on that learns the HUD as part of the world's dynamics.
I ran into this myself. I trained a video predictor on raw Minecraft footage and it spent a suspicious amount of capacity modeling the hotbar. The model believed the inventory bar was part of the physical world. Remove the player's hand from the frame and it would still generate one.
Game2World formalizes the fix. The paper introduces GameUI-Taxonomy, a taxonomy of 21 UI categories with 5,132 verified elements, and GameCleaner, a mask-free model that identifies and removes HUD elements while preserving the scene underneath. No masks needed at inference time, which matters because per-frame mask annotation doesn't scale to millions of gameplay hours.
The numbers are the point. World models trained on UI-free gameplay improve VideoReward by 6.83% over models trained on UI-overlaid data. That gain is the difference between a model that understands game physics and one that's memorized a UI overlay. GameCleaner hits 95.36 AAR on synthetic video, beating the strongest temporal mask baseline by 57.3%, and 80.05 in-the-wild with 99.8% background preservation.
Quick Take: The bottleneck in embodied AI has shifted from model architecture to data acquisition, and the field is splitting between those who find cheaper data and those who generate it.
Egocentric video: the structured middle
EgoStandard sits in the middle tier. 90,000 hours of head-view egocentric video with synchronized 3D hand pose. 80,000 hours with hand pose only, 10,000 more with full-body pose. To put that in perspective: 90,000 hours is over a decade of continuous footage. No lab could collect that with a Franka arm.
The hand-pose annotations are the key. They give you a proxy for action without teleop hardware. Pair that with inverse dynamics and you can recover pseudo-actions from video that was never collected for robot training. This is exactly the trick Veeda's founders describe: using inverse dynamics and latent action models to expand action-conditioned training data beyond what teleoperation can produce.
Access is gated. You have to request permission and state your use case. That tells you something about how the company views the data's value, and it's a reminder that the best egocentric data is still locked behind privacy review. Faces, license plates, and other PII get blurred, but the review process is real.
Teleop data, standardized
The expensive tier is getting cheaper to use, if not cheaper to collect. NVIDIA's Cosmos3-DROID converts the full DROID dataset into LeRobotDataset v3.0 format. 71,907 episodes, 22.4 million frames, 707 GB. Three synchronized camera streams, joint state, torques, cartesian poses, and natural-language task labels. It's released under the OpenMDW 1.1 license, so commercial work is on the table.
The format matters more than it sounds. DROID originally shipped in RLDS format, which is a known pain. My team spent weeks fighting RLDS parsing before switching to LeRobot tooling. The v3.0 standard groups episodes into chunked Parquet and MP4 shards, so you can stream training data instead of loading per-episode files. That's the difference between waiting out an hour-long download and starting training in minutes.
One detail: the conversion splits episodes into success/ and failure/ directories. 57,639 success episodes, 14,268 failures. Most robot datasets quietly drop the failures. For world models, failure trajectories are signal. They show what doesn't work, which is exactly the kind of negative data that interactive learning needs. Also note the conversion isn't a perfect mirror of the original release: the DROID paper reports 76K success episodes and 16K failures. Some episodes didn't survive the format migration.
The four tiers, side by side
By now the pattern is clear. Each tier of the data hierarchy has a dedicated effort to make it usable:
| Tier | Source | Scale | What you get | Main cost |
|---|---|---|---|---|
| Teleop | DROID via Cosmos3-DROID | 71,907 episodes, 22.4M frames | State, action, 3 camera views, language labels | 50 collectors, 12 months |
| Egocentric | EgoStandard | 90,000 hours | Head-view video + 3D hand pose | Access review, PII processing |
| Gameplay | Game2World | 96K synthetic + 1,079 in-the-wild clips | UI-free interaction video | UI removal and verification |
| Generated | Veeda (target) | Unlimited rollouts | Action-conditioned simulation | Compute, drift control, fidelity |
The generated tier is the wildcard, and it's where the biggest bet just landed.
The generative route: Veeda
Veeda AI raised $90M in seed funding, led by Khosla Ventures and Radical Ventures. The team is Sanja Fidler, formerly NVIDIA's AI research VP who ran the spatial intelligence lab, plus Zan Gojcic and Huan Ling. Their thesis: robots can't learn by trial and error in the real world, because it's unsafe, expensive, and slow. So the simulation has to come from world models.
The lineage matters here. Fidler's team went from GameGAN (an AI replacing a game engine) to DriveGAN (trained on autonomous driving video) to latent diffusion video models to Cosmos, and finally to OmniDreams, an action-conditioned generative world model for closed-loop autonomous driving simulation. OmniDreams rolls out minutes of continuous driving in response to actions, and NVIDIA used it to rank driving policies with results consistent with real-data reconstruction.
The shift in the last few years is from "generate a realistic world" to "generate a world you can train policies in." Those are different problems. Realism is necessary but not sufficient. A world model for robot training needs to respond to actions, stay stable over long rollouts, and not let objects vanish or clip through each other.
The 36kr piece reconstructs Veeda's likely stack from the founders' prior work and job postings: action-conditioned video models plus latent dynamics, built on rectified-flow diffusion transformers with block-causal autoregressive hybrids and a causal video tokenizer. The tokenizer's compression ratio directly determines how long the model can keep predicting. To fight drift, they plan to train on the model's own rollouts and use diffusion forcing, where the model is conditioned on its own previous predictions rather than ground truth, which keeps long-horizon generations coherent.
The endgame is a world model that plugs into Isaac Lab and CARLA, closes the loop with a policy, and lets you post-train models like π0 and GR00T N in simulation before touching a real robot. Then you measure the sim2real gap and iterate.
Planning inside the world model you have
LeFlow attacks a different problem: how do you actually use a world model once you have one? The standard approach treats the world model as a black-box simulator and runs iterative trajectory optimization from scratch for every state-goal pair. Every replanning step pays the full optimization cost again, and none of the planning experience carries over.
LeFlow's answer is to amortize planning. A rectified-flow model learns a latent trajectory prior directly in the world model's latent space. Given a current state and a goal, it imagines a future latent path between them. An inverse dynamics decoder turns those latent transitions into action chunks, and the frozen world model verifies each candidate by autoregressive rollout. Planning becomes conditional latent trajectory generation with a fixed budget, instead of an optimizer searching from zero.
The results: consistent success-rate gains across four goal-conditioned pixel-control benchmarks, with an order-of-magnitude reduction in planning time. Practical translation: if your planner took 500ms per replan, LeFlow gets it to roughly 50ms. That's the difference between a robot that visibly pauses to think and one that reacts smoothly.
The deeper argument is the one that matters: latent world models should support reusable planning priors, not just prediction. Once you've learned a world model, you shouldn't throw away the planning experience you could accumulate inside it.
Common pitfalls
Five things trip people up in this area.
Training world models on raw gameplay footage. I made this mistake. The model learns the HUD as part of the world's dynamics, and then your "world model" is generating health bars and minimaps instead of physics. Game2World's 6.83% VideoReward gain shows the cost. Clean the UI first.
Treating egocentric video as if it has action labels. It doesn't. EgoStandard gives you hand pose, which is a great proxy, but you still need inverse dynamics to recover actions, and that's error-prone. Train an action-conditioned model on ego video without accounting for this and you're learning a correlation, not a controller.
Using the wrong data format. DROID's original RLDS format is a known time sink. The LeRobot v3.0 conversion exists because people kept burning weeks on parsing. If you're starting a project today, don't fight the old format.
Evaluating world models on single-step prediction. A model can look great at one-step-ahead prediction and diverge completely after 200 autoregressive steps. Diffusion forcing and self-forcing exist because compounding error is the real problem. Evaluate on rollout length, not just loss.
Replanning from scratch every time. If you're using a world model as a black-box simulator and running iterative optimization for every state-goal pair, you're paying the full cost repeatedly. LeFlow's amortized approach is an order of magnitude faster. If you're building a planner, think about what planning experience you can reuse.
One thing to remember
World models are only as good as the data they're trained on, and right now the data problem is being attacked from four directions at once. The winning stack probably doesn't pick one tier. It cleans gameplay video for pretraining, uses egocentric data to recover actions, standardizes teleop for post-training, and generates the rest in simulation.
What this means for your robot stack
Training video world models on gameplay footage? Run a UI-removal pass first. Game2World's 6.83% VideoReward gain is the cheapest improvement you'll find this year.
If you're building a planner on top of a latent world model, don't treat it as a black-box simulator. Amortize planning with a trajectory prior like LeFlow's, and you get an order-of-magnitude speedup with equal or better success rates.
For startups deciding where to spend effort, the generated-data tier is the one to watch. Veeda's $90M seed is a bet that interactive learning inside world-model simulation replaces a large chunk of real-world data collection. If long-horizon stability keeps improving, the cost curve of embodied data flips within a year or two.