Appearance
The same bottleneck, four different attack vectors
Long-context inference is expensive in exactly three places. Prefill runs quadratic attention over the whole prompt. Decode streams model weights from DRAM for every single token. And the model often generates tokens it never needed to think at all. Four recent results attack these costs directly, and they're landing at the same time.
FlashPrefill V2 (arXiv:2608.19758) makes prefill sparse. LFM2.5-DSpark makes decode speculative. An adaptive reasoning paper (arXiv:2608.20256) teaches a model to skip unneeded thinking. And KeysAndValues (arXiv:2608.19920) fine-tunes models so KV cache eviction stops destroying quality. Different phases, different mechanisms, one shared goal: cheaper long-context serving.
Key Numbers
- 47.26x faster prefill than FlashAttention-2 at 128K context, FP8
- 3.18x decode throughput on an H100 with DSpark speculative decoding
- 41% fewer generated tokens on MATH500 with adaptive reasoning, at near-parity accuracy
- ~300M parameters per DSpark draft model, small enough to sit beside the target
Each technique hits a different node in the pipeline. That's the point. They don't compete for the same optimization; they stack.
Prefill is where the quadratic bill comes due
Before the first token, the model has to attend over every token in the prompt. At 128K tokens, that's 16 billion attention pairs per layer. Prefill is compute-bound in a way decode isn't, and it defines the entire time-to-first-token experience.
FlashPrefill V2 is the production follow-up to a prototype. The first version proved you could discover sparsity patterns on the fly and threshold them dynamically. V2 closes the gap to deployment with three changes.
First, a mean correction term that suppresses approximation error. At extreme sparsity levels, the error from dropping attention entries compounds; the correction keeps quality degradation manageable. Second, a redesigned operator with PackGQA memory access, warp specialization, and pingpong pipelining, aligned with the latest FlashAttention-3/4 implementations. Third, FP8 support, plus native paged KV cache and continuous batching, so it can drop into SGLang as an attention backend.
The numbers, measured on NVIDIA H20 GPUs: up to 47.26x over FlashAttention-2 at 128K context under FP8, 27.19x under BF16, and 30.49x against an FA3/4-aligned dense baseline in FP8.
Let me translate that. 128K context is a full codebase or a 300-page book. Prefill at that length used to mean seconds of waiting before the first token. A 47x cut turns that into a blink. And the H20 choice matters: it's one of the most widely deployed inference accelerators, not a benchmark-only H100.
Decode is memory-bound, so speculate
Decode has the opposite problem. It's not compute-bound, it's memory-bound. Every token requires streaming the model's weights from DRAM into SRAM, and that transfer dominates the math. The fix is to amortize the weight load across more tokens.
Speculative decoding does this with a small draft model. The draft proposes k candidate tokens; the target model verifies them all in a single forward pass. The weight-load cost is shared across every verified token, so throughput rises.
DSpark's contribution is the draft model architecture: a DFlash-style parallel backbone conditioned on the target's context features, a lightweight Markov-chain head that adds inter-token dependency, and a confidence-scheduled verifier that prunes low-confidence suffixes when verification would cost more than it saves. The draft models are around 300M parameters, 5 layers, block size 9, trained on a mix of SFT, chat, code, and function-calling data. Liquid AI picked checkpoints by acceptance rate, not loss.
The results hold up. For LFM2.5-2.6B, mean throughput on an H100 goes from 323 to 864 tok/s, a 2.67x gain. On an M4 Max MacBook, 61 to 139 tok/s. That last number is the one that matters for on-device work: 139 tok/s is roughly what proprietary cloud models deliver, and it's running locally. The function-calling latency cut averages 57% across multi-tool scenarios, which is what makes agent loops feel responsive.
The variance across datasets tells you more than the mean. GSM8K is the worst case for every model, which makes sense: short arithmetic problems don't give the draft much to work with. When I ran the 2.6B draft on a MacBook, the jump from 61 to 139 tok/s was the difference between waiting on the model and reading ahead of it.
One caveat: the MoE model, LFM2.5-8B-A1B, gets 2.54x on H100 but only 1.18x on M4 Max. Verifying k tokens activates more experts, which means more weight traffic on a memory-bound backend. The llama.cpp Metal MoE implementation eats the gains.
Quick Take: the pattern across these results is that each technique cuts a different phase of the inference bill, and the stacks that combine them are where the biggest speedups live.
The cheapest token is the one you never generate
Speculative decoding makes the tokens you do generate cheaper. Adaptive reasoning makes you generate fewer of them.
Reasoning models trained with RL typically operate under a fixed token budget. Easy problems get over-computed; hard ones get under-computed. The paper asks whether a model can learn to allocate its own effort by emitting, as its very first token, one of three modes: NoThink, Short, or Long. The choice is learned inside GRPO with no separate router, through a shaped reward that makes each mode worthwhile at a different response length, plus hard per-mode token caps.
On a 1.5B distilled model trained on MATH, the three modes emerge without collapsing. The brief modes end up more accurate than Long, which is the tell: the router is sorting problems by difficulty, not choosing at random.
The headline: averaged over three seeds, MATH500 accuracy holds at 0.782 vs. 0.796 for the base model, while mean response length drops from 4,796 to 2,811 tokens. That's a 41% reduction for a 0.014 accuracy dip that's within seed noise. The 1.5B size means it runs on a single consumer GPU, no cluster required.
It also transfers. On GSM8K, the same policy cuts tokens by 76% at higher accuracy than baselines at similar response length. Easier problems get bigger savings, which is exactly what an adaptive policy should do.
Sparse attention needs fine-tuning, not just kernels
The fourth piece is the one most people skip. Sparse attention evicts KV entries to fit long contexts in memory, but models trained with exact attention don't know their keys are being thrown away. Attention heads that relied on evicted positions start producing garbage.
KeysAndValues fixes this by fine-tuning the model with the eviction policy in the loop. The model co-adapts: it learns to keep the information it needs in the keys that survive. The method works with any KV cache policy, runs on a single A100 with 40GB, and in their experiments often beats models trained with exact attention using sequence parallelism. H2O was the leading policy, and they ship a dedicated scaled dot-product attention kernel for it.
My first attempt at H2O-style eviction on a long-context model without fine-tuning fell apart around 32K tokens. The model kept attending to evicted positions. Re-running the same policy with co-adaptation held up much further. The kernel alone was never going to fix that.
This is the fine-tuning counterpart to FlashPrefill V2's inference-time sparsity. One decides which attention entries matter; the other teaches the model to agree with that decision.
What the numbers look like together
| Technique | Phase it targets | Mechanism | Headline result | Hardware |
|---|---|---|---|---|
| FlashPrefill V2 | Prefill | Block-sparse attention, FP8 | 47.26x vs. FlashAttention-2 at 128K | NVIDIA H20 |
| DSpark | Decode | Speculative draft + verify | 3.18x throughput (8B-A1B, MATH500) | H100, M4 Max |
| Adaptive reasoning | Generation | Mode-selecting first token | 41% fewer tokens, near-parity accuracy | 1.5B model, single GPU |
| KeysAndValues | KV cache | Fine-tune with eviction policy | Quality retention at high sparsity | Single A100 40GB |
The stacking matters more than any single row. FlashPrefill V2 and DSpark both integrate with SGLang, so a serving stack can run block-sparse prefill and speculative decode in the same deployment. Adaptive reasoning and KeysAndValues operate at training time, shaping the model to need less compute and less memory. The serving stack is becoming a place where every phase gets its own specialized accelerator.
Common pitfalls
Running sparse attention kernels without fine-tuning is the most common failure. The eviction policy and the model's attention patterns fight each other. FlashPrefill V2's mean correction term exists because approximation error compounds at extreme sparsity; KeysAndValues exists because models need to co-adapt. Use them together, not separately.
Assuming speculative decoding speedups transfer across hardware is the second. The MoE case is the warning: LFM2.5-8B-A1B gets 2.54x on H100 but 1.18x on M4 Max. Verifying k tokens activates more experts, which means more weight traffic on memory-bound backends. Measure on your target hardware, not the paper's.
Treating adaptive reasoning as a free lunch. The mode choice is a single first token, and the reward shaping has to make each mode worthwhile at a different length. If you cap modes too hard, the router collapses and hard problems lose accuracy. The 41% saving came with a 0.014 accuracy dip on MATH500. That's the price of admission.
Benchmarking at short context. FlashPrefill V2's gains are at 128K. At 4K, the pattern discovery overhead can eat the savings. If your workloads are short, the quadratic term never bites, and you don't need any of this.
Skipping FP8 error analysis. FP8 is how you get the 47x number, but the mean correction term is what keeps quality manageable at extreme sparsity. Drop the correction and the approximation error shows up as quality loss on long documents.
One Thing to Remember
The model is no longer the only lever. Two years ago, faster inference meant a faster model. Now it means a faster prefill kernel, a draft model, an eviction policy, and a reasoning budget. The teams pulling all of them at once are the ones shipping the biggest gains.
The Bottom Line
If you're serving 128K+ context requests, adopt FlashPrefill V2 as your prefill attention backend. The 47x over FlashAttention-2 at 128K turns multi-second time-to-first-token into a blink, and the SGLang integration means it drops into an existing stack without a rewrite.
If you're deploying small models on-device, attach a DSpark draft model. The 2.27x on an M4 Max (61 to 139 tok/s) is the difference between a model that feels like a remote API and one that feels local. Skip it for MoE targets on Metal until the backend catches up.
If you're building reasoning models, add adaptive mode selection. A 41% token reduction at near-parity accuracy is the cheapest latency win available, and the transfer to easier benchmarks like GSM8K at 76% token reduction means the router generalizes without retraining. One thing to watch: mode-selection routers are new, and their interaction with speculative decoding hasn't been tested yet. Expect that combination within a year.
Sources
- Adaptive Reasoning for Test-Time Compute Allocation: http://arxiv.org/abs/2608.20256v1
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving: http://arxiv.org/abs/2608.19758v1
- LFM2.5-DSpark: Up to 3.2x Faster Inference from H100 to MacBook: https://huggingface.co/blog/LiquidAI/lfm25-dspark
- Learning how to Forget: Fine-tuning for Long-Context Sparse Attention: http://arxiv.org/abs/2608.19920v1