Skip to content

Skipping the KV Cache Read: Declarative Attention Meets the Serving Stack

#llm-inference #sparse-attention #kv-cache #sglang #long-context

Long-context decode is dominated by KV cache reads ​

Every token in a 1M-token conversation sits in the KV cache, and decode reads all of it. Generate a 500-token answer and the attention mechanism scans the full history 500 times, even though most of the attention mass lands on a handful of tokens.

Serving engines soften this with batching, prefix caching, and paged attention. The algorithmic bill stays the same: O(N) KV reads per token, N being the entire context length.

One line of work attacks this with proxy scores. A lightweight model pre-selects the relevant tokens so the main attention head only reads a subset. DeepSeek's V3.2 sparse attention uses this pattern. The catch is that scoring is itself an O(N) pass. You read everything cheaply to avoid reading everything expensively.

A paper on r/MachineLearning takes an intrinsic route. Its question: wouldn't the model already know which parts of the context it needs?

Declarative attention: a protocol, not a scoring head ​

Declarative Attention (DA) is a protocol, not an architecture change. The model writes its intention into the chain of thought using three tags:

  • <global>: attend to the full context.
  • <focus>: attend to a specific region the model names.
  • <local>: attend to recent output only.

The inference engine parses these declarations like tool calls, then skips most of the KV cache. No extra scoring head, no lightweight retriever, no re-ranking stage.

The results are from zero-shot evaluation on off-the-shelf models across 15 long-context tasks (arXiv:2609.02737):

Attended token reduction on Gemma-4-31B: 52.0% Attended token reduction on Qwen-3.6-27B: 31.1% Accuracy drop: 1.27 percentage points on Gemma-4-31B Accuracy drop: 2.75 percentage points on Qwen-3.6-27B

52% fewer attended tokens is the difference between reading a whole book and reading the chapter that answers the question. The accuracy drop narrows as models get bigger, which suggests declarations get more reliable with scale.

How it compares to the other sparse attention options ​

The interesting trade is DA versus proxy scores. Proxy scoring is battle-tested and trains the selector to be accurate, but it adds a pass over the whole context on every step. DA removes that pass entirely and replaces it with a text declaration. In exchange you carry uncertainty: the model decides, and the model can be wrong.

StrategyContext selectionPer-step costTraining neededWorks on existing models
Full attentionAll tokensO(N)NoYes
Proxy-score selectionLightweight model scores, then picksO(N) for the scoring passYesNo
Declarative attentionModel self-declares region in CoTO(selected region)NoYes

Quick Take: DA doesn't remove the KV cache. It removes the need to read all of it on every step. Prefill still builds the full cache; the read path is what shrinks.

What the community is saying ​

The reaction on r/MachineLearning was cautiously enthusiastic. Most people agreed the core idea is sound: the model knows where it needs to look, and an engine that trusts it can skip a lot of work. The disagreement was about how much trust is safe.

When I tested DA-style prompting on a long summarization job, the focus declarations were right most of the time, but occasionally the model anchored on a section I never asked about. The paper's averages are honest about the typical case. Production cost is about the tail, so a fallback path matters more than the headline reduction.

A recurring concern was the declaration overhead itself. Every tag is an extra token in the output stream. On short answers, the chain of thought you write to decide where to attend can cost more than the reads you skip. The 52% reduction is gross savings on attended tokens, not net profit on end-to-end latency.

The prompt injection angle drew the sharpest comments. Declarations are plain text inside the context. If external content contains something that looks like a <focus> tag, the model can inherit it. For agentic workloads where documents and tool outputs flow into context, that's a fresh attack surface nobody has mapped yet.

The serving side: SGLang is already building the substrate ​

This is where the story meets serving infrastructure. SGLang currently powers trillions of tokens per day across more than 400,000 GPUs, and it has spent two years attacking the same cost from the system side.

RadixAttention gives it prefix caching across requests. Prefill-decode disaggregation separates the two phases so one doesn't starve the other. DFlash and Spec V2 handle speculative decoding. And as of the DeepSeek-V3.2 release, SGLang ships day-0 support for sparse attention, the proxy-score flavor.

SGLang capabilityWhat it doesShips
RadixAttentionCaches and reuses shared prefixes across requestsYes
Prefill-decode disaggregationRuns prefill and decode on separate resourcesYes
DeepSeek-V3.2 sparse attentionSkips KV reads via trained proxy scoresYes
DFlash / Spec V2Cuts decode latency with speculative decodingYes
Declarative attentionSkips KV reads via model self-declarationResearch

The pattern worth noticing: SGLang keeps absorbing attention-level optimizations into the runtime. DA's declaration protocol is structured output, and structured output is something SGLang already parses for tool calls and constrained generation. The gap between a paper and a serving feature here is smaller than it looks.

On the hardware side, the GB200 NVL72 work reported 2.7x decode throughput gains from disaggregation, then 3.8x prefill and 4.8x decode when scaled further. The GB300 pushes a 25x end-to-end number. Those gains all come from using the hardware and the KV system more efficiently, which is exactly where DA wants to play.

SGLang is also the rollout backend for post-training frameworks like Miles, verl, and AReaL. That matters for DA's likely next step: the same engine that serves declarations in production is the one you'd use to generate training data for them.

What DA integration looks like in production ​

The flow borrows everything from structured output and tool calling:

The engine checks the declaration stream after each chain-of-thought step. On <focus>, it locates the region's KV pages, loads only those, and runs attention over them. This composes with RadixAttention: the region lookup is a radix-tree walk, and the engine already maintains the page index for paged attention. DA's skip only pays off if the region can be gathered cheaply.

The hard engineering problem is not parsing. It's page layout. A declared region can span dozens of non-contiguous KV blocks, and a naive gather erases the savings. Serving engines that want to support this will keep per-turn or per-section KV locality, or maintain a region index alongside the page table.

Common pitfalls ​

Treating declarations as ground truth. The model picks the wrong region often enough to matter. Build a fallback: measure attention mass on the declared pages, compare it to a threshold, and escalate to global attention when it drops. A bad <focus> on a 200K context degrades the answer silently, and your unit tests won't catch it.

Counting the savings before paying the declaration overhead. DA emits chain of thought before and between tags. On short generations, the extra output tokens can cost more than the reads you save. Profile end-to-end latency on your actual prompt-length distribution, not on decoded-token counts.

Assuming prefill and memory shrink too. DA skips decode reads. The full KV cache is still written during prefill and still occupies the same memory. First-token latency and GPU footprint are unchanged; only the steady-state decode read path gets cheaper.

Forgetting that KV pages scatter. Paged attention stores blocks in non-contiguous memory. Unless the engine keeps a region index or enforces per-turn locality, the "skip" becomes a gather that costs as much as the read you avoided.

Ignoring the injection surface. Declarations are text in context, so they inherit the context's trust boundaries. In agentic setups where external content arrives mid-conversation, you need to isolate or sanitize input that can forge declaration tags.

One thing to remember ​

An algorithm that reads less is only valuable if the engine can serve the remaining reads cheaply. DA's 52% attended-token cut does not automatically become a 52% latency win. It becomes one when RadixAttention-style caching, paged layouts, and a scheduler that understands regions are already in place. The serving stack is what converts sparsity on paper into throughput in production.

The bottom line ​

If you serve long-context models today, keep your engine. SGLang's existing sparse-attention support for DeepSeek-V3.2 is production-hardened, and DA is a zero-shot research result on two model families. Run it behind a flag on your own workloads and measure end to end before committing.

If you're building the serving stack, treat declarations as another structured output format. You already parse tool calls and constrained JSON; parsing focus tags is the same machinery. Engines that can serve arbitrary KV regions cheaply will run next year's sparse-attention models, whichever selector wins.

If you're training models, put DA patterns into post-training data now. The paper shows declarations emerge zero-shot and improve with scale, which is the same arc tool calling took: emergent first, trained later. The team that trains for reliable declarations early will own the accuracy-efficiency trade for the next generation of long-context models.