Skip to content

The Image Editing Stack Just Split Into Three Layers

#image-editing #diffusion-transformers #lora #comfyui #super-resolution #open-source

The image editing stack just split into three layers ​

Google is rolling out a closed image editor this month. The open source side answered with a fully reproducible training recipe, a paper that adapts frozen super-resolution models by touching eight channels per block, and a pile of community LoRA packs. The threads connect: image editing now has layers, and each layer moves at its own speed.

The top layer is the polished product. Google Pics, built on the Nano Banana image model, is rolling out over the coming weeks to Google AI Pro and Ultra subscribers and most Workspace business customers. It does object segmentation with text comments on specific regions, in-image text editing and translation that preserves the font, shared collaboration, and multiple generations from one prompt. It works inside Docs, Slides, and Drive, so you can edit an image without leaving your document. The integration starts in Docs and Slides, with Drive following, and a standalone version lives at pics.new. For a designer who works inside Workspace all day, that removes an entire class of context switching.

None of it is exportable. Nano Banana has no open weights and no public model API outside Workspace, and the editing behavior rides on the subscription. You can use the product. You can't build on it.

The bottom layer moved this month too. LLaDA-Image, a 6B diffusion transformer trained from scratch, released its weights, training code, and detailed recipes. That is a different kind of open than dumping a checkpoint. People can reproduce the pipeline, not just download the result.

In between sits a thin adaptation layer that barely existed two years ago: small LoRA packs for Qwen-Image-Edit and Krea 2 Turbo, plus research like SPARK that steers a frozen DiT model by modulating fewer than ten channels per block. The most interesting work in image editing has moved to that middle layer. The table below maps the layers against what they offer.

LayerRepresentative projectModel accessWhere it runsWhat it's for
Closed productGoogle Pics (Nano Banana)Subscription, no weightsWorkspace apps, pics.newIn-document editing for non-technical users
Open foundationLLaDA-Image, Qwen-ImageWeights, training code, recipesLocal GPUsReproducible research, product backbones
AdaptationLoRA packs, SPARK controllersOpen artifactsComfyUI, HuggingFace SpacesNew behaviors without full fine-tunes
OrchestrationComfyUIOpen sourceLocal, cloud, desktopNode graphs, production pipelines

6B parameters for LLaDA-Image, roughly 12GB of VRAM in bf16, so a 24GB consumer card runs it. 8 channels per stream and block is all SPARK modulates; the SR backbone and VAE stay frozen. 53.53 / 53.38 on Qwen-Image-Bench, the best open-source scores on the English and Chinese tracks. 2 to 4 sampling steps for the distilled Turbo variant, which puts interactive editing latency within reach.

LLaDA-Image: a training recipe you can repeat ​

Pay close attention to what LLaDA-Image (arXiv 2609.03796) does before the text-to-image stage. The team builds a visual prior through image-only pre-training and mid-training, then runs a generation pipeline of 220M samples, most of them synthetic. Only about 98M of those samples are real images. A frozen vision-language module built on the LLaDA2.0-Mini diffusion language model backbone handles the text side, and the model learns fine-grained editing instructions without the paired-image-text-heavy diet earlier work assumed was mandatory.

The training details matter if you want to reproduce any of it. The DiT uses parameter-free RMSNorm throughout, meaning the norm layers carry no learned scale and shift weights, which reduces memory. The optimizer is Muon rather than Adam, and that is not a cosmetic choice. Muon replaces Adam's per-parameter momentum with orthogonalized updates, changing both memory use and convergence behavior. If you are used to training DiTs with AdamW, don't eyeball these hyperparameters. Follow the recipe.

The results justify the effort. LLaDA-Image scores 53.53 on the English track and 53.38 on the Chinese track of Qwen-Image-Bench, the best numbers any open-source model has posted on either. For practical work, the distilled Turbo version is the more interesting release. Generating in 2 to 4 sampling steps puts a 6B model in the same speed class as the Krea 2 Turbo workflows people already run in ComfyUI.

SPARK: eight channels beat a fine-tune ​

On the super-resolution side, SPARK (arXiv 2609.03813) starts from a simple observation: activations in pretrained DiT SR backbones are dominated by a small number of massive channels, and controlled interventions show those channels strongly affect reconstruction quality. So instead of fine-tuning the network or bolting on another adapter, SPARK trains a lightweight input-conditioned controller that predicts bounded per-channel scale and shift values for the dominant channels. An online activation-ranking procedure finds which channels matter for the specific backbone being adapted. The SR backbone and VAE never change.

The controller conditions on the low-resolution VAE latent, so it learns rules like "when the input looks like this, turn up these channels." It works: consistent gains on DIV2K, RealSR, and DRealSR across three DiT-based SR backbones, with only eight channels modulated per stream and block.

The controlled comparisons matter. The paper verifies that the gains are not explained by parameter budget or by mere access to the selected channels. The information the controller carries does the work. That is a different adaptation axis than LoRA. LoRA changes weights across thousands of dimensions. SPARK changes activations in eight. If this generalizes beyond super-resolution, the cheapest way to specialize a DiT may be a small network that learns which channels matter, not a weight update across the whole model.

Quick Take: the fastest-moving part of image editing right now is the adaptation layer between frozen backbones and final output, not the base models themselves.

The LoRA layer is where the community tests ideas first ​

The community has been voting. Three HuggingFace Spaces show where demand sits: QWEN_EDIT_IMAGE, a copy of the Qwen-Image-Edit 2511 LoRA pack, at 176 likes; the Rapid-AIO experimental LoRA bundle for Qwen-Image-Edit at 101; and the Krea 2 Turbo image-to-image space at 97. The numbers are small by model-download standards, but they climbed quickly, and they point the same direction: people want editing behavior, fast iteration, and LoRAs they can swap without retraining.

When I loaded the 2511 Qwen-Image-Edit LoRAs into ComfyUI, the change from the base model showed up within a few generations. Instruction following tightened, and text rendering, historically the weak point of these models, got noticeably better. The Rapid-AIO experimental pack is rougher, the name says experimental, but it bundles a wider set of styles and behaviors into one download. The Krea 2 Turbo space makes a different trade: speed first, iterate on composition, accept some precision loss.

My main frustration is base-model mismatch. LoRA packs train against one checkpoint revision and behave differently on another. I swapped in an older Qwen-Image checkpoint once and the edits silently degraded. No error. Just worse outputs. Every LoRA pack should state its base model hash up front.

ComfyUI turns the stack into something you can run ​

ComfyUI's job is unglamorous: it makes every layer usable at once. The node graph natively supports the open models that matter right now, including Qwen Image, Flux.2, Krea 2, Hunyuan Image 2.1, Ideogram 4, and Z-Image. Partner nodes reach closed models like Nano Banana, Seedance, and Hunyuan3D. The image-editing list is just as long: Flux Kontext, Flux.2 Klein, Qwen Image Edit, HiDream E1.1, OmniGen2, MageFlow Edit, and LongCat Image Edit. When a model drops on Friday, a workflow usually exists by Monday. ComfyUI runs a weekly release cycle, with a stable core branch shipping a major version roughly every two weeks.

The execution engine is what makes it production-grade. Asynchronous queueing, partial graph re-execution so only changed nodes rerun, smart VRAM management and model offloading, support for quantized models, and workflow recovery from generated images. Drag a PNG back onto the canvas and you get the full graph including seeds. App Mode exposes a complex workflow through a simple UI, and a local API pushes workflows into production pipelines. Offline use is a first-class option: --disable-api-nodes cuts off the optional paid API nodes. It runs on NVIDIA, AMD, Intel, Apple Silicon, and Ascend, which matters now that editing models land on every platform.

The layer boundaries are concrete. A Google Pics user never touches the model. A ComfyUI user switches from Qwen-Image-Edit to Flux Kontext by rewiring a few nodes. Model releases, training recipes, and adaptation tricks improve independently, and that independence is why the stack moves as fast as it does.

Common pitfalls ​

  • Don't mix LoRA packs across base revisions. The 2511 Qwen-Image-Edit LoRAs target a specific checkpoint. On an older base they load without error and quietly produce worse edits. Save the base model hash alongside the LoRA metadata.
  • Don't run ComfyUI on a stale PyTorch. The project treats torch 2.7 as the floor and recommends the latest stable release, and on NVIDIA 20-series and above you want a cu130-or-newer build. Old torch shows up as cryptic OOMs and wrong colors, not as a version error.
  • Treat SPARK-style controllers as trained adapters, not plug-and-play modules. The dominant channels are identified per backbone through the online ranking procedure, and a controller trained on one DiT SR model will not transfer to another.
  • Don't assume a fully open recipe is cheap to reproduce. LLaDA-Image's 220M-sample pipeline with a 6B DiT is serious compute. The openness buys reproducibility and research leverage, not a budget training run. Plan accordingly.
  • Don't build a product core on Google Pics. The Nano Banana behavior is subscription-gated and tied to Workspace. It is a capable editor, not a model you can integrate. For programmatic access and pipeline control, the open stack is the stack.

One thing to remember ​

The base model is no longer the differentiator it was a year ago. Google Pics, LLaDA-Image, Qwen-Image-Edit LoRAs, and ComfyUI all build on the same generation architecture family. The gaps between them come from adaptation and orchestration: which channels you modulate, which LoRAs you stack, which nodes you wire together. That is why the small artifacts, an eight-channel controller and a few hundred megabytes of LoRA weights, are where the progress is happening.

The bottom line ​

If you are building a product around image editing, anchor on open weights such as LLaDA-Image or Qwen-Image inside ComfyUI rather than on a closed subscription model. You get the same class of editing behavior with full control over inputs, outputs, and inference cost.

If you are latency- or VRAM-constrained, use the distilled paths: LLaDA-Image-Turbo at 2 to 4 steps, or the Krea 2 Turbo workflows. The speed win outweighs the fidelity loss when you are iterating interactively.

Watch SPARK's activation-modulation line. Adjusting eight channels per block is competitive with fine-tuning on super-resolution today. If the idea generalizes beyond SR, full model fine-tunes for editing will start to look overpriced within a year.