Appearance
The two bets in visual generation
Open the arXiv listings and the Hugging Face Spaces feed from the past few weeks and you'll see the field arguing with itself. One camp pushes everything into a single network that understands images, reasons about them, and generates new ones in one forward pass. The other camp wires small, specialized models together into graphs and treats the graph itself as the product.
Both camps released something concrete recently. SenseNova-U1.5 is an 8B mixture-of-transformers model that handles understanding and generation natively, no vision encoder, no VAE. LoopVAE rethinks the visual tokenizer, the layer most pipelines treat as plumbing. Workflow1111 rebuilt AUTOMATIC1111's feature set as a 73-node Gradio graph where every output is a REST endpoint.
They're solving different problems, and the gap between them is where the interesting engineering decisions now live.
Key numbers
- 8B MoT parameters: SenseNova-U1.5 runs inference on a single workstation GPU, no cluster needed.
- 4K native resolution: it handles poster-sized images without tiling or chunking.
- 0.28 rFID on ImageNet-256: LoopVAE hits this with 29M parameters, about 65% fewer than the 84M reference VAEs.
- 65% to 90%: detector-based object accuracy across a 3.5-day training run of a 210M DiT.
- 22 of 73 nodes: Workflow1111 runs these entirely in-process with no network call, so two-thirds of the graph works offline.
SenseNova-U1.5: one network for everything
The unified argument is simple: if a model can read and write pixels, the boundary between understanding and generation disappears. SenseNova-U1.5 is built around that claim. It's encoder-free, so there's no frozen CLIP or SigLIP front-end, and VAE-free, so image tokens live in the same space as text tokens instead of passing through a separate latent compressor. The practical consequence: you can feed it an image, have it reason about what's wrong, and ask it to emit a corrected image in the same context. At 8B parameters with mixture-of-transformers routing, this runs on a single GPU for inference. That number matters more than any benchmark score.
The post-training recipe is the part to copy. The team strengthened the visual interface with spatially coherent patch reconstruction, then trained specialized experts for visual aesthetics, bilingual text rendering, infographic generation, and image editing. They consolidated all of that through multi-expert on-policy distillation. The experts stay quietly available behind one interface without model surgery or weight merging. They plan to open-source the training code: supervised fine-tuning, reinforcement learning, and the distillation itself.
The result that matters most: despite limited exposure to structured formats in the generation data, the model generalizes to long, complex, structured visual instructions. That's evidence that multimodal understanding transfers to visual planning. It's the strongest argument yet for the native unified approach.
Quick Take: The highest-leverage work in visual generation right now is at two interfaces: the attention maps inside the DiT and the graph that wires models together. Base model scale is no longer the main lever.
What a single-GPU training run teaches you
I spent 3.5 days training a 210M text-to-image DiT from scratch on one RTX PRO 6000, using 4.2M images at 256². The goal was to understand the recipe end to end, and the full write-up is worth keeping around. Three measurements came out of that run that I hadn't seen stated plainly anywhere else.
First, the learned null attention slots become the sink. The model uses 16 register tokens in the image stream plus 2 learned key/value slots appended to every cross-attention. At mid-noise in a middle block, those 2 slots absorb roughly 90% of the cross-attention mass. The EOS token, the usual sink in cross-attention models, drops to about 4%. Content words keep a few percent each, sharply focused on their objects. Register vectors grow to 4 to 13 times the norm of image tokens by the middle blocks.
Second, treat the flow-matching loss as a health signal. It says nothing about output quality. The loss moved 0.805 to 0.754 over the whole run while held-out FID went from 33.7 to 27.0, FD-DINOv2 from 570 to 218, and detector-based object accuracy from 65% to 90%. Most of the loss at high noise is the irreducible variance of the velocity target. Training and held-out loss stayed equal to the third decimal for 24 epochs. If you're watching the loss curve to decide when to stop, you're watching the wrong thing.
Third, the training-time timestep shift beats doubling the steps. On 2,456 held-out prompts with the final weights: 20 steps with shift 2.8 gives FID 27.0, 50 steps gives 26.6, 8 steps gives 28.4, and 20 steps without the shift gives 27.3 with FD-DINOv2 degrading from 218 to 228.
Attention maps as control signals
DIAL pushes the same thread further for multi-subject video generation. The authors found that certain attention blocks in Diffusion Transformers naturally form an Intrinsic Spatial Grounding Map (ISGM) that precisely locates reference subjects. In low-noise stages, they use the ISGM to steer attention, giving precise control over fidelity strength at inference time without retraining. In high-noise stages, the same maps build preference pairs at no extra cost for reinforcement learning, anchoring attention to reference subjects and mitigating semantic drift. It consistently beats baselines on the OpenS2V-Eval benchmark for identity consistency.
Put the two results side by side. The single-GPU run shows that what attention attends to is often a sink, not a semantic target. DIAL shows that the residual spatial signal in those maps is still precise enough to guide generation and to generate rewards. Both treat DiT internals as a surface you can measure and manipulate, which is a different mindset from the black-box API approach of the SD1.5 era.
What the community is doing with these models reinforces the point. The Spaces that get traction are the ones you can duplicate and rewire. The Krea-2-Turbo_I2I Space has 135 likes. A Wan 2.2 I2V 14B Space with custom LoRA support sits at 81. When I pulled these up, the pattern was clear: people aren't waiting for the perfect unified checkpoint. They're forking the latest image-to-video model, attaching their own LoRAs, and shipping. The LoRA hooks and the ability to repoint a model node matter more than the base model's benchmark table.
Tokenizers and data: the quiet layer
LoopVAE takes a different route to efficiency: share parameters across scales instead of adding blocks. Hierarchical visual tokenizers normally give each spatial scale its own processing blocks. LoopVAE reuses a scale- and loop-conditioned core within and across scales, keeping only the resolution-changing transitions independent. A four-block core executes 28 block applications per encoder or decoder pass. That's heavy reuse, and it pays off: 29M parameters reach 0.28 rFID and 32.54 dB PSNR on ImageNet-256 with a roughly 30-epoch two-stage training budget, about 65% fewer parameters than the 84M reference.
The honest caveat sits in the runtime profiling: fewer stored weights mean more arithmetic and longer runtime. Parameter sharing trades memory for FLOPs. If you're shipping on an edge device where weights are the constraint, that's a good deal. If you're GPU-bound in a training cluster, it's less obvious.
The plankton paper is the same lesson in a different costume. Automated plankton imaging produces severely long-tailed datasets, and the rare taxa of greatest ecological interest have too few images to train or evaluate classifiers. The fix: generate synthetic plankton conditioned on taxonomy. A CLIP encoder is adapted on a large plankton corpus with a ranked contrastive objective extended to deep, ragged taxonomies, then frozen to condition a parameter-efficient diffusion transformer. The pattern transfers beyond plankton: adapt an encoder on a domain hierarchy, freeze it, condition a small DiT on it, and you have a data-augmentation pipeline for any long-tailed recognition problem.
The connection to LoopVAE: both make the parts around the big model cheaper and more reusable. That's the engineering work that doesn't make headlines but makes everything else possible.
The composable counterpoint: Workflow1111
Workflow1111 is the clearest statement of the modular position. It rebuilt most of AUTOMATIC1111's stable-diffusion-webui as a single graph: eleven media pipelines built from seventy-three nodes. Text-to-image, hi-resolution fix, image-to-image, prompt-matrix grids, VLM-based prompt interrogation, detection-to-inpaint masks, ControlNet-style annotators, background removal, PNG Info, and image-to-video are all paths through one canvas.
The design details are where this gets interesting. Of 36 operator nodes, 32 are plain Python fn nodes, and 22 of those run entirely in-process with no network call. Roughly two-thirds of the canvas keeps working if you lose your connection. The only remote dependencies are the model calls themselves. A workflow is debuggable in proportion to how much of it runs locally.
Two decisions make this more than a ComfyUI clone. First, every output node becomes a typed REST endpoint with no hand-written routes. Workflow1111 exposes nine: /image, /edited_image, /generated_prompt, /recovered_prompt, /detected_objects, /x_y_grid, /upscaled_local, /annotator_map, /png_info. Second, those same endpoints are MCP tools. Launch with mcp_server=True and every output node shows up as a tool an agent can call from Claude Code or Cursor. An agent can generate an image, read a prompt back out of it, and run detection as steps in a larger task, with no glue code.
The graph also runs entirely on other people's hardware through Inference Providers and Hub Spaces, so you can build it without a GPU of your own. But the fn node abstraction means you can equally bind a local checkpoint and drive it from the same canvas. The graph doesn't care where compute lives.
Side by side: unified vs. surgical
| Decision point | Native unified (SenseNova-U1.5) | Composable graph (Workflow1111) |
|---|---|---|
| Architecture | Single encoder-free, VAE-free network | Node graph of specialized models |
| Integration | Shared token space, end-to-end | Frozen backbones joined by operators |
| Control | Prompt, expert routing, distillation | Rewire nodes, per-pipeline parameters |
| Hardware | 8B MoT, one GPU for inference | Serverless per-node, or local fn nodes |
| Extensibility | Retrain or distill new experts | Add a node; endpoints generate themselves |
| Failure mode | Shared failure surface | Failures isolate to a single node |
| Best fit | Understanding + generation in one context | Deterministic, inspectable media pipelines |
The honest framing: these aren't competing products in most deployments. They optimize different bottlenecks. SenseNova collapses the pipeline for tasks where you want one context that reasons about pixels and then changes them. Workflow1111 optimizes the harness: deterministic stages, testable functions, auto-generated APIs.
Common Pitfalls
Watching the flow-matching loss as a quality metric. The single-GPU run saw loss move 0.805 to 0.754 while FID improved by 20%. At high noise the loss is dominated by irreducible velocity variance. Track FID, FD-DINOv2, or a detector metric instead.
Skipping the timestep shift. The shift rule from SD3/RAE, √(32·32·32/4096) for the 32-channel FLUX.2 latent, beat doubling the sampling steps. 20 steps with shift beat 50 steps without it. If you're training a rectified flow model, treat the shift as a hyperparameter, not an afterthought.
Reading attention maps as if content words are the carriers. With 16 register tokens and 2 learned slots, roughly 90% of cross-attention mass lands on the slots at mid-noise. Attention-based interpretability and control need to account for sinks, or your importance scores measure the wrong thing.
Truncating recurrent tokenizer loops. LoopVAE found that completing the trained recurrence improves reconstruction, and truncation exposes output-range errors rather than graceful degradation. If you reuse a core across scales, run the full loop at evaluation.
Freezing backbones without a communication path. For joint image-layout generation, InterIL keeps both image and layout diffusion backbones frozen but learns a communication module between them. Freeze the bases and just concatenate outputs, and you get two good single-modality priors that never coordinate. The interaction is where the joint distribution lives.
Assuming every node in a workflow needs a model call. In Workflow1111, 22 of 32 fn nodes run in-process with no network. Design the graph so the only remote dependencies are actual inference, and two-thirds of your pipeline stays debuggable offline.
One thing to remember
The field is not converging on a single architecture. Unified models are getting better at long, structured, interleaved tasks, and composable graphs are getting better at orchestration, APIs, and agent integration. The strongest systems a year from now will probably combine both: a native unified model doing the heavy cognitive work, wrapped in a graph that exposes its outputs as endpoints and lets agents rewire around it. The winning stack lives in the seam between the two.
Choosing between unified and surgical approaches
If you're building an application that needs understanding and generation in the same context, edit-by-description with reasoning, or interleaved captioning and synthesis, adopt a native unified model like SenseNova-U1.5. A single token space eliminates the serialization overhead and identity drift you get from chaining separate models, and 8B MoT parameters keep inference on one GPU.
If you're constrained by infrastructure or need deterministic stages you can inspect and test individually, detection to mask to inpaint, skip the monolithic route and wire specialized models into a graph like Workflow1111. Half the graph runs without a GPU or network, every output is already a REST endpoint, and the same endpoints become MCP tools with one flag.
One thing to watch: attention-based control is turning DiT internals into a steering surface. DIAL's ISGM gives inference-time fidelity control without retraining, and the sink measurements tell you exactly where to look inside the attention map. Expect retraining-free control to become a default feature of image and video pipelines within two quarters.