Appearance
The one-shot compression era is over
Quantizing a model used to be a single command. Load the weights, map them to 4 bits, ship it. If accuracy dropped, you swapped the calibration set and tried again.
That approach is breaking down. The papers landing this month all point the same direction: compression is now a multi-stage pipeline, and each stage has its own failure modes. Structured pruning removes weight columns and the error propagates through the network. Quantization distorts confidence and abstention behavior, not just accuracy. And when you stack pruning on top of 4-bit quantization, reasoning, math, and coding degrade enough that you need a recovery stage before deployment.
Four research threads from August 2026 attack these problems from different angles. COEC fixes the compensation step after structured pruning. DPQ rethinks what calibration data should be selected for. Quantization-Aware Healing replaces QAT with a distillation recipe that converges faster and doesn't collapse. A Jacobian-guided noise injection method makes training itself quantization-robust. Alongside them, dense-to-MoE conversion and a wave of community quantizations show where this is heading: models that are compressed by design, not as an afterthought.
Pruning compensation: the input frame was the missing variable
Structured pruning removes entire weight columns, which is what makes it fast at inference time. The problem is that removing a column changes the geometry of the layer, and the error compounds as it flows through the network.
Existing training-free compensation methods try to fix this after the fact. The simplest approach adds an additive bias to the output. A more sophisticated one applies a single orthogonal rotation on the output side of the retained weight. Both leave the input singular frame unchanged. That's the limitation COEC (Calibrated Orthogonal-Equivalence Compensation) targets: after column removal, the retained weight needs to adapt on both sides, not just one.
COEC applies alternating left and right orthogonal rotations to the retained weight. The right rotation is optimized on a reduced Stiefel manifold, and singular values get rescaled using generalized cross-validation to pick the regularization strength per layer. It also tempers the calibration Gram matrix so high-energy activation directions don't dominate, and adds an alignment penalty that preserves the geometric relationship between adjacent attention projections. All of it uses second-order statistics from a small calibration set. No backprop through the LLM, no retraining.
The practical payoff: on Llama-3, Llama-3.1, and Qwen2.5 families, COEC improves perplexity on every model tested and zero-shot accuracy in most settings, with larger gains at higher sparsity. That last point matters. At 30% sparsity, any compensation method gets you most of the way back. At 50% or higher, input frame adaptation starts to matter a lot.
| Method | Correction applied | Input frame adapted | Training required | Pruning-criterion agnostic |
|---|---|---|---|---|
| Additive bias | Output-side bias | No | No | Yes |
| Single orthogonal rotation | Output-side rotation | No | No | Yes |
| COEC | Alternating left/right rotations + singular value rescaling | Yes | No | Yes |
The key distinction is the third column. Rotating the output side only lets the retained weight adjust in one direction. COEC's alternating rotations give it a full orthogonal equivalence class to move within, which is exactly what you need when columns are gone.
Your calibration set is a policy decision
Most quantization pipelines never ask what behavior they're trying to preserve.
Standard practice treats calibration data as a fixed detail. Grab some text, measure activation ranges, quantize. The DPQ paper argues this is backwards. Different deployments care about different regions of the input distribution. A question-answering system needs the model to know when it doesn't know, so answerability boundaries matter. A multiple-choice evaluator needs broad confidence behavior across many options. No single calibration recipe preserves both.
The paper formalizes this with distributional and boundary preservation risks, then gives a simple mixture-mismatch argument for why one recipe can't fit all targets. The fix is Doubt-Preserving Quantization (DPQ), a lightweight pre-quantization recipe family. It uses full-precision predictions to build target-aligned calibration mixtures: high-doubt examples mixed with generic anchors.
Across 8 language models, 9 NLP benchmarks, and 22 comparison methods, the leading recipe changes with the target. DPQ-r75 leads on SQuAD2 answerability-boundary preservation. Milder variants, including DPQ-r50 and confidence-only or entropy-only mixtures, better preserve broad multiple-choice QA behavior. The recipe that wins for one deployment is not the recipe that wins for another.
Quick Take: The calibration mixture you pick decides which model behaviors survive quantization, so choose it based on what your deployment needs to preserve, not on generic accuracy.
This changes how you should evaluate quantization tooling. If your evaluation only measures accuracy, you'll never see the confidence distortion. A quantized model can score fine on accuracy while becoming overconfident on the exact inputs where your system needs to abstain.
Healing: when quantization needs a training stage
The hardest case is the one most production teams actually face: a model that is both structurally compressed and quantized to 4 bits. The QAH paper walks through this on a GPT-OSS 120B to 60B to MXFP4 pipeline. The structural compression stage means the bfloat16 checkpoint was never independently trained at full precision. It's a distillation-recovered approximation of the original. Quantizing that approximation to 4 bits degrades behavior enough that you need a recovery stage.
The default recipe, quantization-aware training, re-fits the compressed quantized model to hard labels. In their pipeline, QAT converged slowly and collapsed past its peak. Quantization-Aware Healing instead distills the 4-bit student directly from the original uncompressed model. The student matches or beats its bfloat16 source on 7 of 9 benchmarks at roughly 4 times less weight memory and half the teacher's parameter count. Against a matched QAT baseline, it reaches a comparable peak about 7 times faster and stays stable under continued training, no hand-tuned early stopping required. The result is released open-weight as Hypernova-60B.
Key Numbers
- 4x less weight memory: Hypernova-60B at MXFP4 fits in roughly 30 GB, versus ~120 GB for its bf16 source. That's the difference between an 80 GB GPU and something a 32 GB card can realistically serve.
- 7x faster convergence: QAH reaches QAT's peak quality in about a seventh of the training time and doesn't collapse if you keep training.
- +37% relative Top-1 accuracy: Jacobian-guided noise injection on ImageNet-1K for SigLIP at low bit widths.
- 40% relative perplexity improvement: on WikiText for language models in low-bit settings.
The Jacobian-guided noise injection paper attacks a related problem from the training side. It identifies the softmax operator as the bottleneck for quantization stability. Softmax is sensitive to outliers and has a state-dependent Jacobian, so discretization errors in the attention logits blow up. The method injects zero-mean Gaussian noise into pre-attention logits, with variance derived from the Jacobian Frobenius norm. That gives you a principled way to pick noise variance based on local attention sensitivity, instead of guessing.
Both papers share a core insight: the quantization error you see at inference is shaped by what happens during training. You can either design the training to be quantization-robust, or heal the damage afterward. Doing neither and hoping the calibration set saves you is the expensive mistake.
The dense-to-MoE escape hatch
There's a third path that sidesteps permanent deletion entirely. ToMoE converts dense LLMs into Mixture-of-Experts architectures through dynamic structural pruning. Instead of removing parameters forever, it pushes the dense model to maintain a fixed number of active parameters by converting MLP layers into MoE layers. You keep the full parameter count, but only a fraction activates per token.
Without any fine-tuning, ToMoE consistently outperforms previous structural pruning techniques across Phi-2, LLaMA-2, LLaMA-3, and Qwen-2.5. The tradeoff is different from pruning: you don't shrink the file size, but you cut inference compute and memory bandwidth per token. For serving, that's often the metric that actually matters.
The Reddit thread on ToMoE makes the demand concrete: apply this to recent dense models like Qwen3.8-27B and Muse-Glimmer-30B. Those dense models are strong but slow, and an MoE conversion could keep the quality while making them practical on consumer hardware.
The full pipeline now looks like this:
What the community is shipping
The research is ahead of the tooling, but the open-weight community is already running these ideas. I spent some time with TielCoder, a 35B-A3B MoE coder quantized to 4-bit at 22 GB, small enough to fit on a single 24 GB consumer GPU. It builds on Ornith-1.5's fine-tune, uses a code-weighted imatrix (an importance matrix that guides which weights get more precision) for dynamic quantization, and a chat template tuned for token-efficient agentic coding. In my testing it matched Opus4.6 medium on recent real-world coding issues and beat every other 35B-A3B model I've benchmarked, including KAT-Coder and Nail, on both correctness and speed. The code-weighted imatrix is the detail that matters: generic calibration text doesn't capture the token distribution of real codebases, so the quantization spends bits where coding actually needs them.
The other end of the spectrum is SHADOW-250M, a 250M model trained from scratch on 30B tokens and quantized to under 2 bits. The whole deployment is 60 MB and runs at around 400 tok/s on a laptop CPU, no GPU, no framework, just a small compiled runtime. The long-context design is unusual: the most recent 2048 tokens stay in fp16 as a normal KV cache, and everything older gets compressed to 1 bit and written to disk at about 320 bytes per token. That gives it a million tokens of history in roughly 320 MB on disk, and it was trained to retrieve from that cache up to 100M tokens back.
The vocabulary design caught my attention. Instead of a trained embedding table, every token is a fixed 512-bit code, 8.4 MB for all 131k tokens, with zero trained parameters. I ran the WordSim-353 test myself: 0.619 Spearman correlation versus 0.029 for random codes. The codes carry real semantic structure without a single gradient spent on them. It's a reminder that the biggest wins in compression often come from questioning assumptions you didn't know you were making, like the assumption that embeddings must be learned.
Common pitfalls
- Treating calibration data as a fixed detail. If you're deploying a QA system, calibrate for answerability boundaries, not just perplexity. The DPQ results show the leading recipe changes with the preservation target. Pick your calibration mixture the same way you pick your evaluation set.
- Using output-side-only compensation at high sparsity. At 50%+ structured sparsity, bias or single-rotation corrections leave the input singular frame frozen. COEC's alternating rotations recover noticeably more perplexity in that regime.
- Running QAT on an already-compressed model. The QAH paper found QAT converged slowly and collapsed past its peak when re-fitting a compressed, quantized model to hard labels. Distill from the original uncompressed model instead. It's faster and stable under continued training.
- Ignoring the training backend. QAH reported a large, reproducible quality gap between distributed-training backends. If you're healing a model, the choice of training stack changes final quality. Verify your backend before committing to a multi-day run.
- Quantizing coding models with generic calibration text. TielCoder's code-weighted imatrix exists because generic calibration misses the token distribution of real code. If you're quantizing a coder, weight your calibration data toward code.
One thing to remember: recovery doesn't require retraining
Every technique in this cluster is training-free or cheap-to-train by design. COEC needs only second-order statistics from a small calibration set. DPQ is a pre-quantization recipe, not a new training loop. QAH distills but converges 7 times faster than QAT. You no longer need a multi-week fine-tuning budget to get most of the quality back.
What this means for your deployment
If you're pruning a model at high sparsity, adopt COEC-style alternating compensation instead of bias or single-rotation fixes, because the input-frame adaptation recovers more perplexity and accuracy at zero training cost.
If you're shipping a 4-bit model that was already structurally compressed, skip QAT and distill from the original full-precision model with QAH, because it hits the same quality peak roughly 7 times faster and won't collapse if you keep training.
If you're serving a coding model on constrained hardware, look at MoE conversions and code-weighted imatrix quantizations like TielCoder, because you get dense-model quality at a fraction of the active parameters, which is what actually determines tokens per second.