Skip to content

World Models and Video Generation Are One Problem Now

#world-models #video-generation #jepa #latent-transition #diffusion-models #interactive-video

Two Fields Colliding on the Same Failure Mode ​

World modelers and video generators started with different questions. World models ask what state a scene will be in next. Video generators ask what the next pixels should look like. Those questions are becoming one question, because both sides keep hitting the same wall: things that should stay stable drift over time.

A train enters a tunnel. A world model knows the train still exists and will come out the other side. A frame-by-frame video generator has to guess whether to draw it again, and often guesses wrong. A character gets introduced as "the woman in the red coat," then later prompts call her "she," and her face shifts every shot. An e-commerce product spins on screen and grows a sixth finger or the wrong price tag.

The evidence arrives from both directions at once. This month's research batch includes a JEPA encoder that grows only as the task demands, a latent transition architecture built around temporal evolution instead of token mixing, an object-conditioned human motion forecaster, an image-to-scene generator grounded in 3D proxies, and a training-free fix for character consistency in visual storytelling. On the product side, Decart sells real-time interactive video through AWS and was valued at $4B in May. Kling raised $3B at an $18B valuation in July, and is watching its most replaceable revenue stream erode.

What used to be a rendering problem is now a consistency problem. Identity, geometry, and state have to survive over time, and every serious player is converging on that constraint from a different starting point.

Encoders That Grow When the Task Does ​

JEPA (joint-embedding predictive architecture) world models use a Vision Transformer encoder to map observations into a representation space where prediction happens. Standard practice is to pick a fixed-size encoder and hope it covers every task. That fails on both ends. Simple tasks waste capacity across redundant attention heads. Hard tasks run out of abstraction.

Successive Capacity Growth works the other way. The encoder starts minimal: one attention head, two transformer layers, 283K parameters total. A task-agnostic test-and-verify loop proposes an expansion. Width additions (more attention heads) buy low-level semantic capacity. Depth additions (more transformer blocks) buy higher-order abstraction. Because each trial uses function-preserving expansion, the bigger encoder begins with exactly the behavior of the smaller one. The loop evaluates prediction loss and rolls back any expansion that doesn't help.

A regularizer, the Sketched Isotropic Gaussian Regularizer, keeps the learned semantic dimensions statistically independent and aligned with the predictive objective. That matters as the architecture grows, because larger encoders collapse into redundant subspaces if nothing pushes them apart.

The results should make you question the fixed-size default. On a 60-dimensional multi-object dynamics task, SCG naturally triggered depth expansion and improved prediction loss by 20.3% over the fixed small baseline. Compared with scaling a fixed model to the large configuration, it was 56 times more parameter-efficient. On a 2D navigation task, a single width expansion beat the fixed large model by 23%. The loop produced zero false-positive expansions and preserved the original function bit-exactly.

283K parameters is a laptop-GPU encoder, not a cluster job. The efficiency claim matters more than it sounds: JEPA training is data-hungry, and an encoder that doesn't need to be big cuts the batch sizes and data volumes required to cover it.

Key numbers: 283K starting parameters (1 head, 2 layers). 20.3% prediction-loss improvement over the fixed small baseline. 56x parameter efficiency vs. the fixed large model. 23% gain from one width expansion on navigation. 0 false-positive expansions, bit-exact function preservation.

The Transition Is the Model ​

Robot policies increasingly rely on World Action Models, which predict how task-relevant scene state evolves under interaction. The trend is to run those predictions in latent space. It's cheaper than rendering pixels and keeps control-relevant information. But there's a mismatch hiding in plain sight: most WAMs realize the latent transition with a Transformer, and a Transformer's inductive structure is organized around token interaction, not temporal evolution. It's a general-purpose mixer doing a specialist's job.

LEON treats transition realization as its own architectural choice, separate from the predictive representation and separate from how prediction couples to policy. It learns an observable space, then propagates state through a context-modulated operator, grounded in the controlled Koopman generator view of dynamical systems. Context-dependent variation is organized around a shared evolution-operator structure. A second path, additive forcing, covers what operators can't express: discrete events like object pick-ups, contact, or external interventions that break the learned linear flow. Without that forcing term, every perturbation gets flattened into the shared operator. With it, the operator stays stable and the residual has somewhere to go.

That split is the design. The operator carries the dynamics that repeat across contexts. The forcing term carries the changes that don't fit. When the agent hits a new object or a collision, the model doesn't reshuffle its whole evolution structure. It applies a forcing correction.

Across two WAM formulations that integrate latent prediction into the policy differently, LEON improved closed-loop performance and robustness, and held up even when the transition module was fully replaced. For a roboticist, the last point is the practical one: you can swap the transition predictor without retraining the policy coupling around it.

The Pixel Path Learns to Track State ​

The pixel path is where most of the money went, and three papers this month show the same correction: diffusion models need conditioning that encodes state, not just appearance.

Object-Conditioned Social Diffusion folds motion history, all-to-all social interaction, and object cues into one framework. The object-conditioning mechanism modulates denoising at every timestep. That per-step modulation is the detail that matters. The model isn't told once to look at the stove. It's reminded at every rung of the denoising ladder, which is how human-object reasoning survives into the generated futures. OCSD tops the Humans in Kitchens and HOI-M3 leaderboards, cutting two-second path error by 121.5mm (31.3%) and 130.5mm (33.2%). A third less error at the two-second horizon is the difference between a robot that anticipates and one that reacts.

SpatialCrafter takes on the other consistency problem: camera motion. Turning one image into an explorable scene requires the world behind the frame to exist. The framework splits generation into a global 3D proxy stage and an appearance stage. A Point-anchored Sparse Structure Flow module predicts a spatially aligned, geometrically consistent proxy. Then a pre-trained video diffusion model, reframed as a Generative Deferred Refiner, paints high-frequency photorealistic detail onto the proxy-defined geometry. Parallel Geometry Injection and Proxy-Aware Corruption training keep the VDM from fighting the proxy. The team built a 115K-scene hybrid dataset because nothing suitable existed. The win: stability under rapid camera motion and extreme viewpoint changes, where naive video diffusion drifts into hallucination.

Sidecar attacks the quiet problem in free-form storytelling. A character is fully described when first introduced, then referenced by "the girl" or "he" in later prompts. Those type-level mentions drop the identity-defining semantics, and generation drifts. Sidecar keeps the initial entity description and injects the missing semantics into later prompt embeddings. No training, no architecture change, negligible overhead, and it improves prompt-image alignment and character consistency across SDXL and FLUX baselines.

The first time I ran a WAN 2.1 LoRA for a character pipeline, training was the easy part. Dataset prep ate the week: cut detection, caption hygiene, reference frames. And after all that, identity still drifted by the fourth or fifth second of generated video. The failure mode was identity, not rendering. The community is deep in the same fight; a single published Wan LoRA tutorial repo pulled 31K downloads last month. Sidecar's design is the one I'd borrow for my own pipeline, because it treats consistency as an injection problem instead of a training problem.

PaperProblem it attacksMechanismMeasured result
SCGFixed-size JEPA encoders waste or starve capacityFunction-preserving width/depth growth with rollback + SIGReg20.3% loss gain, 56x parameter efficiency, zero false expansions
LEONTransformer transitions don't model temporal evolutionKoopman-style operator + additive forcing in a learned observable spaceBetter closed-loop robustness across two WAM formulations
OCSDHuman motion ignores objects and group dynamicsObject-conditioned denoising at every timestep + social encoder121.5mm and 130.5mm path error cuts (31.3%, 33.2%)
SpatialCrafterSingle-image scenes drift under camera motionGlobal 3D proxy (PaSS Flow) + VDM as deferred refinerStable under extreme viewpoints; 115K-scene dataset
SidecarLater prompts lose identity semanticsTraining-free semantic injection into prompt embeddingsConsistency gains across SDXL and FLUX baselines

Quick Take: Consistency over time, not rendering quality, is the constraint that now binds world models and video generation together.

Real-Time Interactive Video Is the Product Test ​

Product companies are testing the same convergence in the market, and the numbers are big. Kling AI reported Q2 revenue past 850M CNY, up more than 200% year over year. It raised nearly $3B in July at an $18B post-money valuation. Decart raised about $300M in May at a $4B valuation, with NVIDIA among its investors, and Anthropic reportedly floated a $6B acquisition in August. Both companies are being priced as if the bet already paid off.

The product difference from last generation is interactivity. Sora, Seedance, Veo: you type a prompt, you get a finished clip, you can't touch it. Real-time interactive video generates from the current frame and your input, live. Decart's Lucy, distributed through AWS on a token basis, watches your room and changes it. Virtual try-on that tracks your body as you turn. That's a new interaction mode, and it's the first one on this stack to convert into revenue.

There's still a gap between revenue and a killer app. The most visible Lucy demo shows a face distorted into a swollen shape. It looks cool for a second, and then you remember a phone filter did this years ago. If the category lands as a fancier filter, the technology and the valuation will start to drift apart.

Kling's financials show how fragile its footing is.

Nearly 80% of Q1 revenue came from overseas, and over 60% of total revenue ran through the API channel. Eight of the ten largest API customers in the first five months were overseas. In April, three overseas super-customers reduced their calls because of competition. At that concentration, a rival's model upgrade is a threat to the P&L before it's a threat to the product.

Three Routes, Three Bets ​

Under the interactivity umbrella, three technical routes are racing. The world model route simulates hidden state. The pixel route paints frames. The code route keeps a scene list. They differ in exactly the dimension the research predicts: how each one maintains what should stay true.

RouteWhat it maintainsStrength in practiceWeakness that shows upWho's building it
World modelHidden state: the train still exists in the tunnelObject persistence, reasoning under occlusionExpensive compute and dataDecart
Pixel / VDMNothing beyond the current frameVisual quality, immersionObject counts, text, prices driftKling, Seedance, Veo, Wan, Vivix
Code / DLMA scene list: the train is marked "occluded"Accuracy, editability, low costVisual polish lags the pixel routeSigmaZ AI (Tap8)

The dark horse is the code route, because of what powers it: diffusion language models. An LLM generates one token at a time, each token waiting on the last. A DLM lays down a rough full draft, then refines all positions in parallel, which is why it can be faster and why it extends to longer sequences. The approach lived mostly in small models until this year. Ant Group's LlaDA 2.0 pushed a DLM past 100B parameters. ByteDance open-sourced Cola DLM in May, and Google DeepMind released Diffusion Gemma weights in June. In real-time video, the code route builds exactly on this: a scene graph that grows like a live webpage, where every item is a record you can drag, rotate, or delete. Open weights are in the race too. The 14B Wan 2.2 image-to-video model sits on Hugging Face now, and it needs a serious GPU, or quantization, to run at full quality.

SigmaZ built a consumer product first, and the customers redirected them. I've met the merchant they describe, or someone close enough. He uploads product detail pages and asked whether his goods could rotate 360 degrees in a generated video. The tools he'd tried capped at 15 seconds, cost too much, and drew the product wrong. He offered to pay three or four times as much for a version that got the object right. That's the accuracy wedge, and it's why SigmaZ's Tap8 positions as a platform for marketing, sales, and how-to content instead of entertainment.

The founders skew young and the capital is comfortable with that. Decart's CEO is 27. SigmaZ's co-founders were born in 2003 and 1995. Vivix's founder hit research director at SenseTime by 26. The money is betting that young teams can run side bets the big labs won't commit to, and that some of those side bets become the main event.

The Business Problem Nobody Solved ​

Kling can't win forever by chasing Seedance into model and API price wars, argues a 36Kr analysis. ByteDance is moving downstream into the film and TV chain. Jimeng launched a film incubation program, and a Seedance Studio production platform is rumored. Kling positioned itself as a professional film tool, but aside from the fluid effects in 太平年, the analysis counts few concrete moves, and the model updates have slowed.

The alternative: an alliance with a long-video platform. Platforms hold IP, scripts, directors, production management, and distribution. They also have a structural problem that AIGC speaks directly to: content costs are rigid while membership, ads, and user growth are under pressure. AIGC is one of the few levers that expands content supply without expanding cost proportionally. Kling brings the model layer, API stability, and control tools. The platform brings production feedback that prompt-scale data can't: why a shot works, how a scene keeps character and emotional continuity, how a script becomes a finished cut.

That feedback loop is the same semantics-over-time problem Sidecar works on, at industrial scale. It's also a data moat. API customers multi-home and reallocate calls by benchmark. A production partner doesn't. Kling has the model but no production system; platforms have the production system but can't carry frontier model R&D alone. The clock is running. If Kling keeps spending its energy on overseas API volume and leaderboard position, it could miss the moment the domestic AI film production system consolidates.

Common Pitfalls ​

The papers and the products agree about where the failures concentrate. Five to avoid:

Chasing resolution when the failure is consistency. If characters drift or objects change count, more pixels won't fix it. The fix is semantic: Sidecar-style injection, an explicit scene list, or an object-conditioning layer that modulates every denoising step.

Pre-allocating capacity at maximum size. SCG shows a minimal encoder with function-preserving expansion and rollback beats both fixed-small and fixed-large baselines. If your data is scarce or your compute is metered, start small and let the loss function prove that growth is needed.

Using token-mixing transformers as temporal transitions. In latent world models, attention is a poor default for evolution. LEON's operator-plus-forcing transition beats the transformer baseline on closed-loop robustness and can be swapped in without retraining the policy coupling.

Trusting a pure VDM with exact content. Prices, SKU counts, finger counts, text. If the output must be factually right, a frame-painting model will eventually fail it. That's the e-commerce argument for the code route, and it's why OCSD conditions on objects at every timestep instead of once.

*Concentrating revenue on multi-h