Appearance
Four papers, one bottleneck
The arithmetic is brutal. Every token you add to an LLM's context grows attention compute quadratically and the KV cache linearly, and you pay both before the model writes a single answer. Video makes it worse: a short clip becomes hundreds or thousands of visual tokens, and each one flows through the same attention machinery as a word.
Four recent papers attack that cost from different points in the pipeline. None of them is a new flagship model. Two wrap existing models with cheap decision layers. One removes a layer from inside the model. The last one is a map of everyone else's attempts.
The useful way to read them together is as a pipeline, because each paper operates on a different stage. ConvMem shortens the reasoning path: instead of walking a document sequentially, it summarizes segments in parallel and combines the summaries hierarchically. VIP-Router shortens the visual token path: instead of applying one pruning strategy to every image, it picks the strategy per input. RiLM shortens the parameter path: it deletes the output matrix and turns decoding into geometry. The survey covers the video pipeline end to end, which is where the cost curve is steepest.
Key numbers
- 26.9% relative accuracy gain when a router picks a vision token pruning strategy per input instead of using one fixed strategy (VIP-Router).
- 54.2 vs 113.0 WikiText-2 validation perplexity for the hyperbolic small LM versus the best tied recurrent baseline. Roughly 2x.
- 0.017% of backbone parameters is the full cost of the router. The selection logic is nearly free.
- One third of a small model's capacity at d=128 with a 2000-token vocabulary sits in the output projection. RiLM removes it entirely.
ConvMem: turn linear reading into a convolution
The obvious way to handle a document longer than your context window is to read it in segments and maintain a running summary in fixed-size memory. That's MemAgent's approach, and it works, but it has two problems. The reading is sequential, so latency grows with document length. And the memory policy is trained with RL, which is expensive and tends to overfit: software agents pick up parametric priors that serve them in-distribution and quietly fail on out-of-distribution tasks.
ConvMem takes the same building block and rearranges it. Instead of a linear chain of memory updates, you get a hierarchy. The LLM, prompted with the user's query, acts as a convolutional kernel: it slides over text segments, summarizes each one, then summarizes the summaries. The reasoning path collapses from a linear chain into a logarithmic tree.
Three design details carry the weight. Configurable strides control how aggressively the kernel skips ahead to capture evidence at different granularities. Skip connections propagate evidence across levels of the hierarchy, which keeps errors from accumulating as summaries get summarized. Multi-kernel convolution runs several query-conditioned kernels side by side, so a complex query decomposes into disentangled semantic channels instead of forcing one summary to hold everything.
The practical payoff is multi-hop QA, where the evidence is scattered across the document and the question determines where to look. On RULER-HotpotQA and RULER-2WikiMultiHopQA, ConvMem beats training-free baselines and avoids the out-of-distribution drop that plagues RL-trained memory agents.
Here's what sold me. I've watched RL-trained memory policies look great on their eval split and fall apart the moment the document distribution shifts. The parametric overfitting the paper calls out is the same failure I've seen in production. Moving the policy from learned weights to the prompt means the kernel is only as good as your query phrasing, but it also means nothing in the model memorized the benchmark.
Quick Take: the common thread here isn't a new architecture family. It's deciding where the cost should live. ConvMem and VIP-Router leave your weights untouched. RiLM removes a layer inside the model. The survey tells you which existing knobs are worth turning. Match the fix to your actual bottleneck.
VIP-Router: the best pruning strategy isn't one strategy
Vision token pruning exists because MLLMs drown in visual tokens. Hundreds or thousands per image, each paying the same attention tax as a word. The standard approach is to pick the pruning strategy with the best average accuracy and apply it everywhere. VIP-Router's authors found the flaw in that logic: average accuracy hides per-sample complementarity. The average-best strategy wins overall, but on a significant fraction of individual samples, a different strategy is superior. You can't see that in the eval average.
VIP-Router is a lightweight router that reads low-cost visual and textual features and picks the pruning strategy it predicts will work for this specific input, at the requested pruning level. If pruning looks unfavorable, it keeps the option of running full-token inference. The router adds trainable parameters equivalent to 0.017% of the backbone, which is the whole point. You're not paying for an ensemble. The router runs once, picks one strategy, and you pay for one pruned forward pass.
On VTC-Bench Group A, the pruning-sensitive perception benchmark suite, the router beats the best fixed strategy at every reduction ratio: 26.9% relative improvement in average accuracy, and 22.0% relative improvement in utility after accounting for the tokens actually spent. It also transfers across MLLM backbones and shows up on unseen benchmarks without modification.
The idea generalizes past pruning. If strategy selection can be routed per input based on cheap features, the same trick applies to frame sampling rates, chunk sizes, and connector policies. Someone is going to productize this.
RiLM: delete the output matrix, decode with geometry
Sub-million-parameter models matter for edge deployment and domain adaptation, and they have a structural waste problem. At embedding width d=128 with a 2000-token vocabulary, a two-layer LSTM or Transformer spends roughly one third of its parameters on the output matrix that projects hidden states to vocab logits. That matrix does one job, and it eats a third of a tiny model's budget.
RiLM removes that layer entirely. Context unfolds as a trajectory on a Riemannian manifold, and next-token probabilities come from the squared geodesic distance between the current state and the vocabulary embeddings. The same embedding map serves input and output. Decoding is geometry.
Results: HypRiLM, the Poincaré-ball instantiation, reaches 54.2 ± 0.2 validation perplexity on WikiText-2 with ~290k parameters. That's small enough to run comfortably on embedded-class hardware, no cloud GPU needed. Flat RiLM, the Euclidean version, gets 87.6 ± 0.6. Tied and matched LSTM, Transformer, and SSM controls sit at 113 to 147; the best recurrent baseline is the SSM at 113.0 ± 3.8. That's roughly a 2x gap in perplexity.
Two numbers put this in perspective. A third of capacity reclaimed means that for a 300k-parameter model, you get back ~100k parameters of representational budget to spend on the composition map instead of vocabulary projection. And the 2x perplexity gap is the difference between a small model that produces mostly coherent text and one that drifts into confusion.
The catch is trainability. Naive hyperbolic recurrence collapses: states slide toward the boundary of the Poincaré ball and training diverges. RiLM needs Möbius stabilization to restore stable training. If you reimplement this from the equations, you will hit the collapse before you hit the results. Budget for it.
The scope disclaimer matters. The claims are controlled small-model comparisons, not full-vocabulary state of the art. But for the edge segment where models live under a million parameters, the output matrix is where the budget leaks, and removing it changes the trade-off.
The video efficiency map you'd have to build yourself
The survey covers VideoLLMs: systems that couple video representations with pretrained LLMs and condition generation on a text prompt. Their performance on captioning, QA, retrieval, and temporal grounding comes at a cost that grows with frame count and context length, which kills real-time and mobile deployment. The survey organizes inference-efficiency mechanisms by the pipeline stage they act on: frame sampling, modality encoding, connector-level token reduction, and LLM prefilling and decoding.
The most useful thing here is the honesty about evidence. The survey separates accuracy-cost comparisons that share a host model and input protocol from heterogeneous cross-paper evidence. I've been burned comparing FLOPs numbers from papers that used different encoders, different frame rates, and different prompts. Those numbers are directional at best. The shared-protocol numbers are the only ones worth treating as measurements.
| Approach | Where it acts | Training cost | What you pay | Headline result |
|---|---|---|---|---|
| ConvMem | Context assembly | None, prompt-only kernel | Multiple LLM calls, parallelizable | Beats training-free baselines on RULER multi-hop QA |
| VIP-Router | Vision token pruning | 0.017% of backbone parameters | One router forward per input | +26.9% relative accuracy vs best fixed strategy |
| RiLM | Output head | Full small-model training | Removes W_out, needs Möbius stabilization | 54.2 PPL on WikiText-2, ~2x over tied SSM |
| Efficiency survey | Entire video pipeline | None | None | Stage-by-stage map with shared-host-model evidence |
Two gaps stand out. Audiovisual efficiency is thin: most work treats audio tokens as an afterthought, and audio can dominate token count in real footage. And evaluation is not standardized, which is exactly why the survey had to separate shared-protocol results from the rest. The repository is at github.com/momentslab/awesome-efficient-videollm.
Common pitfalls
Fixed-ratio pruning everywhere. If you hard-code a token budget for every image, you're baking in the assumption VIP-Router exists to break. Images with dense, task-relevant detail get gutted; simple images waste tokens. The 26.9% relative gain comes from per-input selection, not from tuning the ratio.
Treating parallelism as free compute. ConvMem shortens the reasoning path from a chain to a tree, which cuts wall-clock latency when you have cores to spare. But you still pay for each kernel invocation across the hierarchy. The win is in latency and error accumulation, not total FLOPs. If you're single-GPU and latency-bound, count the calls before you commit.
Reimplementing hyperbolic decoding without stabilization. The naive recurrence collapses to the boundary of the ball and training diverges. Möbius stabilization is not a tweak; it's the difference between a trainable model and a NaN loss. If your hyperbolic LM won't train, check the geometry before you blame the loss.
Trusting cross-paper efficiency numbers. The survey's shared-protocol distinction exists because cross-paper FLOPs and latency claims are inconsistent. Comparing a FLOPs figure from one paper against a latency figure from another is how you design a system that looks great on a spreadsheet and falls over in deployment.
Forgetting that training-free means prompt-dependent. In ConvMem the LLM is the kernel and the query is the kernel configuration. A vague query gives you a vague kernel. Prompt engineering is now part of the architecture.
One thing to remember: these four works sit at different levels, and the level matters. The survey is a map, not a method; reading it won't cut your latency. ConvMem and VIP-Router bolt onto an existing model and change how it's driven. RiLM changes the model itself. If you pick the wrong level, you'll optimize the wrong stage, and your bottleneck will move instead of shrinking.
The bottom line
If you're doing long-document multi-hop QA and can't stomach RL training, adopt hierarchical convolution. A prompted LLM as a convolutional kernel over segments is training-free, parallel, and sidesteps the out-of-distribution overfitting that RL-trained memory agents suffer.
If you're shipping a multimodal or video model and you prune visual tokens, don't ship one fixed strategy. Add a per-input router; the 26.9% relative accuracy gain over the best fixed strategy comes from sample complementarity that your eval average hides, and you keep full-token inference as a fallback.
If you're building sub-million-parameter models for constrained devices, RiLM's geodesic decoding is the approach to watch. Reclaiming the output-matrix budget is the biggest structural win available at that scale, and the 2x perplexity improvement over tied baselines justifies the stabilization work.