Appearance
Why shrinking the model stops working
Every deployment conversation starts the same way: "the model is too big, can we compress it?" So we prune, quantize, distill, and the state dict gets smaller. Then accuracy erodes, the hardware bill moves sideways, and someone spends two weeks debugging a fused kernel for an 8-bit attention path. Deployment costs track compute and memory bandwidth, not parameter count.
Five recent papers take a different route. Instead of making the model smaller, they skip the work that doesn't change the answer. A training recipe that drops whole layers. An inference-time router that drops low-contribution experts. A RAG compressor that shrinks the context by 16x. A spiking architecture that fires at most once per neuron. A ViT compression stack that cuts 327 MB down to 6 MB. They sit at very different maturity levels, and one of them contains a result that quietly undermines its own elaborate method.
The common thread: all five skip work that doesn't change the answer rather than shrinking the model. Where that applies depends on where your bottleneck sits.
2,400+ layer dropout training runs across models from 271M to 8.2B parameters. 50% of expert slots skipped at top performance with ACE on Qwen3.6-35B-A3B. 16x context compression and 4x-24x inference acceleration with DEX-Comp. 54.5x model size reduction, 327.42 MB to 6.01 MB, at matched accuracy for the ViT stack. 1.5B parameters, the first TTFS-based spiking LLM at that scale.
Layer dropout: a training recipe that keeps paying off
The most FLOPs get spent during training, not serving. Stochastic depth, or layer dropout, randomly drops whole transformer layers during pre-training. It worked well in smaller vision and language models, then quietly disappeared from LLM training because some reports said it hurt accuracy. Nobody had quantified that degradation or bothered to fix the recipe.
The layer dropout study does both. With the right layer distribution, time schedule, and optimizer settings, models trained with layer dropout reach lower validation loss at the same training FLOPs. On a fixed step budget, that's up to 25% of training FLOPs saved. A three-month pre-training run becomes about nine weeks. Those findings come from more than 2,400 experiments spanning 271M to 8.2B parameter models and datasets up to 160B tokens, all run on Cerebras CS-3 systems.
The inference-side benefit is bigger. Because layers were trained under random dropping, the model learns to produce good predictions from subsets of its own layers. You gain early exit, intermediate-layer skipping, and self-speculative decoding. The paper reports up to a 1.5x inference speedup with negligible accuracy loss. A model generating 30 tokens per second moves to 45. No quantization, no kernel rewrites, no new serving infrastructure.
ACE: skipping MoE experts without calibration data
MoE models route every token through the same number of expert slots, whether the token needs them or not. Existing expert-skipping methods lean on router confidence, calibration data, or extra training, none of which reliably measures actual expert contribution. ACE is training-free, calibration-free, and checkpoint-preserving, built from two complementary estimates:
- Global Spectral Proxy (GSP), which estimates global transformation capacity from the coupled gate, up, and down projections together with RMSNorm scaling.
- Router-Conditioned Refinement (RCR), which builds expert-specific direction prototypes from centered router weights and evaluates expert responses along those routing-preferred directions.
During inference, ACE combines both estimates with the runtime router gates and skips an expert slot only when both views identify it as low-contribution. The top-1 expert is always retained. All expert statistics are computed offline, so the online path is table lookups and scalar arithmetic, nothing per-token on the GPU.
The results on Qwen3.6-35B-A3B at a 50% skipping ratio: 7.96% lower WikiText-2 perplexity and 4.15 percentage points higher average downstream accuracy than the strongest competing method. Skipping half the expert compute makes a 35B-total, 3B-active model noticeably cheaper to serve, and the quality gain over other skipping methods suggests the dual-view decision rule is filtering out experts that were hurting the output, not just ones sitting idle. The advantage grows under more aggressive skipping, which is exactly the regime where every competing method collapses.
Quick Take: In all five papers, the win comes from skipping work that doesn't change the output, and layer dropout is the cheapest to adopt because it costs nothing at deployment time.
The five approaches side by side
| Approach | Savings mechanism | Training required | Reported result | Accuracy impact |
|---|---|---|---|---|
| Layer dropout | Random layer drops during pre-training; early exit at inference | Recipe change only | 25% fewer training FLOPs; 1.5x inference speedup | Lower loss at same FLOPs; negligible accuracy loss |
| ACE expert skipping | Token-adaptive MoE expert skipping | None, offline statistics | 50% of expert slots skipped | Better perplexity and accuracy than static baselines |
| DEX-Comp | Soft context compression for RAG | Two-stage RL fine-tune | 16x context compression; 4x-24x inference speedup | Matches or exceeds uncompressed RAG |
| TTFS spiking LLM | One spike per neuron per window | End-to-end SNN training | First TTFS spiking LLM at 1.5B params | Matches ANN on NLU; gap on perplexity |
| ViT compression stack | Hessian-guided pruning + quantization + distillation | Multi-stage pipeline | 54.5x size cut, 327.42 MB to 6.01 MB | 95.13% accuracy, matches FP32 baseline |
Reading down the table, the pattern is clear. The techniques with the smallest deployment footprint, layer dropout and ACE, are the ones you can ship tomorrow. The ones with the largest potential, spiking and context compression, need real training investment and still carry accuracy caveats.
DEX-Comp: compressing the context instead of the model
RAG systems pay prefill cost proportional to everything they retrieve, and most of what they retrieve is filler. Soft context compression encodes each document into a much shorter embedding sequence. The obvious training approach is distillation from the uncompressed RAG system. DEX-Comp points out the flaw: a compression model trained only on a teacher's outputs can never exceed the teacher, because nothing in the training signal asks it to.
DEX-Comp runs two stages. Pure Distillation warm-starts the model on the uncompressed RAG's correct responses only. Hard Exploration then runs reinforcement learning on queries the uncompressed RAG fails, forcing the model to explore computation patterns that work well for compressed representations. The RL stage is the part that can beat the teacher, because it trains on cases where the teacher has nothing to teach.
Results across five open-domain QA benchmarks at retrieval depths from top-5 to top-30: 16x context compression and 4x-24x inference acceleration, with accuracy comparable to or exceeding the uncompressed RAG baseline. A 4,000-token retrieved bundle becomes 250 tokens. That changes the hardware you need: a 7B model handling long RAG queries can move from a pool of accelerators to a single one. The 4x-24x spread shows how much the speedup depends on retrieval depth and how much context you were willing to chop.
The spiking frontier: one spike per neuron
The strangest entry in this batch is the spiking LLM. Time-to-first-spike (TTFS) coding produces at most one spike per neuron within a time window, which gives extremely low firing rates and, in principle, energy-efficient event-driven computation. The blocker was structural: TTFS networks were restricted to specific layer types, and core LLM blocks like layer normalization and matrix multiplication didn't fit.
The paper introduces a reference-based strategy to encode the four core components: embedding layers, layer normalization, attention operations, and dropout. That yields a fully TTFS-based spiking transformer trained end-to-end, scaled to 1.5B parameters, the first TTFS spiking LLM at that size. On BERT and GPT-2 scale models, natural language understanding and common-sense reasoning are comparable to ANN counterparts. Language modeling perplexity still shows a clear gap.
The energy estimate needs scrutiny. The paper reports a spike-count proxy under an established cost model, not a measurement on neuromorphic hardware. So the efficiency claim is arithmetic, not demonstrated silicon. I'd wait for measurements on actual neuromorphic chips before planning any hardware purchase around it.
The ViT case study: when compression doesn't beat a small student
The most production-flavored paper is the ViT compression framework for in-field plant disease detection. It combines Hessian-Balanced Adaptive Block Pruning (H-BAC) guided by second-order sensitivity, then quantization, then attention-based knowledge distillation, staged into a sequential pipeline. On a chilli dataset with an out-of-distribution test split that holds out entire villages and devices, the compressed ViT reaches 95.13 +/- 2.32% accuracy, matching the 95.13% FP32 baseline, across size reductions of 74-98%. The full stack lands at 6.01 MB INT8, a 54.5x cut from 327.42 MB. That size fits microcontroller-class hardware sitting in a field, which is the entire point.
But the more honest result is the direct comparison they included. A student model trained directly at the same final size, with no pruning and no distillation, hits 94.87% accuracy at the same 6.01 MB. That's within noise of the elaborate pipeline. The compression components, on this dataset, don't justify their added cost. When I look at this result, the lesson is to start with a small architecture and train it directly before committing weeks to a compression pipeline. Compression earns its keep when you have a trained large model you can't retrain, or an accuracy gap only the teacher's knowledge can close. Neither was the case here.
Common pitfalls
Compressing first and measuring second is the most expensive mistake. The ViT result shows exactly what happens when you assume an elaborate pipeline beats a plain small model. Run the direct baseline where the small architecture is trained from scratch; the authors needed one comparison to reveal the stack wasn't pulling its weight.
Treating layer dropout as a drop-in default also fails. The gains depend on the layer distribution, the time schedule, and optimizer adjustments, and the paper had to rediscover those after dropout fell out of fashion. Drop it into an existing recipe unchanged and you may reproduce the old accuracy-degradation reports.
Router confidence is not expert contribution. That's the flaw ACE exists to fix. Router probabilities reflect routing preference, not how much an expert changes the representation, which is why ACE requires both the spectral proxy and the direction-based refinement to agree before skipping a slot. Use a single view and you're back to the baselines it beats.
Distilling only on successful RAG outputs is a warm start, not a recipe. DEX-Comp's Pure Distillation stage is deliberately limited to correct responses, but the gains come from the RL stage on failing queries. Skip that and the compressor inherits the teacher's blind spots at 16x shorter context. My team has hit this exact pattern with a distilled compressor that faithfully reproduced the teacher's failures, only faster.
Reading spike-count proxies as energy measurements is wishful thinking. The 1.5B spiking LLM's efficiency numbers come from a cost model over spike counts, not from hardware. Until someone runs it on neuromorphic silicon, treat the energy claims as directional.
One thing to remember
Every one of these papers compares against a baseline that wasn't trying hard enough. The ViT stack lost to a well-trained small student. ACE beats other skipping methods, not exact routing. DEX-Comp's win over uncompressed RAG is a genuinely high bar, but it needs a two-stage RL fine-tune most teams will find heavy. Read the baselines before you adopt the method.
The Bottom Line
If you're pre-training or fine-tuning an LLM on a fixed budget, adopt layer dropout now, because you save up to 25% of training FLOPs and later gain a 1.5x inference speedup through early exit and self-speculative decoding.
If you're serving a MoE model and can't afford calibration data or fine-tuning, implement ACE-style expert skipping, because the offline statistics and dual-view decision rule let you drop half the expert calls while beating every static baseline on quality.
If you're building an edge vision model, resist the compression instinct and train a small architecture directly first; reach for pruning and distillation only when the small student can't close a specific accuracy gap. One thing to watch: within a year, expect at least one major serving framework to ship training-free expert skipping and context compression as standard flags.