Appearance
The real cost of cheap inference
Quantization damage doesn't live where you think. Mobile memory pressure will evict your KV cache mid-conversation. The cheapest model in a six-way race can win on every axis while the most expensive one bills you for silence. These are the three traps of cutting inference cost, and they share a shape: the trouble moves somewhere your existing tools can't see it.
Key numbers
- 21-52 points: what global finer quantization granularity gained over locally repairing the "most recoverable" layers, on every compatible model tested.
- 2.1-5.5x: the Time-to-First-Token reduction mzCache got on-device by restoring evicted memory concurrently instead of blocking.
- 17x: the spread between the fastest and slowest model's time to first token in a six-model race (533 ms to 9.3 s).
- 487 of 512: thinking tokens Qwen3.5-397B spent before its budget ran out, producing zero characters of answer.
- ~12 tok/s: warm decode for a 105 GB mixture-of-experts model streamed from SSD on a 48 GB Mac.
Four results landed in the same window: a quantization damage study, an on-device memory manager, a six-model latency race, and a streaming MoE engine. All four converge on one practical message. You can't tell where the damage is, where the memory went, or which model is fast, without building measurement that looks at the right thing.
Quantization damage is diffuse, not concentrated
Post-training quantization is the standard way to halve serving cost, but its accuracy cost is uneven, and teams usually tune it per model by feel. This study replaced the guessing with a causal mixed-precision intervention. Quantize everything to 4-bit, then raise each layer to 8-bit one at a time and measure how much accuracy comes back. That's the ground truth. Run it across 9 open-weight models from 4 architecture families, and test three cheap hypotheses against it: damage sits in task circuits, in the layers where the model computes, or in weight statistics like outlier scales and norms.
None of them predicted which layers would recover precision.
Instead, recovery is diffuse. For 8 of 9 models, restoring 75% of the accuracy gap takes roughly half the layers. The lone exception is Qwen3-8B, where the damage concentrates sharply. If you're about to spend a precision budget protecting "important" layers, this is the finding to sit with. Your guess is no better than the three hypotheses the paper falsified.
Global granularity beats local repair
At a matched precision budget, the comparison stops being subtle. Spending the whole budget on finer quantization granularity, group-128 instead of per-tensor, beat locally repairing the most recoverable layers on all 8 group-128-compatible models. The margin: 21-52 accuracy points. Even Qwen3-8B, the one model with sharply concentrated damage, does better when the budget goes global. The 9th model, OpenLLaMA, is excluded because its width rules out group-128 quantization.
| Question | What the intervention showed |
|---|---|
| Does damage sit in task circuits? | No. Where the model computes doesn't predict which layers restore precision. |
| Does damage sit in weight statistics? | No. Outliers, norms, and scales don't identify recoverable layers. |
| Does causal intervention find it? | Yes, but recovery is mostly diffuse: ~half the layers restore 75% of the gap in 8 of 9 models. |
| Is peak recovery architecture-specific? | Within a family, yes. Across families, no. |
Two secondary findings matter for practice. The residual is budget-limited: 8-bit weights were near-lossless across RTN, GPTQ, and AWQ in their evaluation, so if you're losing real accuracy at 8-bit, suspect the calibration set or the eval, not the bit width. And the broader lesson: a signal that correlates with quantization damage doesn't necessarily identify where restoring precision helps. That has to be tested with causal intervention.
Quick Take: Every layer of the inference stack has a place where intuition lies, and the fix is always measurement with the right eyes: causal intervention for precision, payload checks for serving, and real worker configs for benchmarks.
Mobile memory pressure: eviction is the problem
On-device LLM inference has a problem that doesn't exist on a server: the OS owns the memory. Users switch apps, pressure spikes, and the kernel evicts the model weights and KV cache. The next inference request either reads everything back from slow storage or recomputes the entire KV cache. On a phone, that's the difference between a warm answer and a spinning loader.
mzCache treats eviction as the design constraint instead of an edge case. It partitions LLM memory into fine-grained shared buffers so the system can evict partially, not whole-model. Eviction uses hybrid swap and backward-out policies chosen for low-latency restoration from any state. The key trick: mobile SoCs have unified memory, so the CPU can restore evicted pages while the GPU keeps running inference. Zero-wait GPU inference with concurrent CPU-side restoration.
Implemented on llama.cpp and deployed as an Android app, mzCache measured 2.1-5.5x lower Time-to-First-Token than storage-backed partial offload in real multitasking scenarios. For a mobile assistant, 2x is the difference between feeling instant and feeling broken after every app switch.
Six models, one race
Separately, a developer got tired of picking a model based on which thread they skimmed last. They built a 390-line Python tool that fires one prompt at six models at once, streams the responses side by side, and shows time-to-first-token and cost per run under each column. The integration is two lines. DigitalOcean's inference endpoint speaks the OpenAI API, so one client and a model string swaps between Llama, DeepSeek, Mistral, Qwen, and OpenAI's open-weight gpt-oss models.
Two things went wrong before any numbers existed. The catalog call returns 72 models and the account can reach six; there's no availability field, so the first run shipped with a dead default model, and the 403 showed up in its own column while the other five streams kept going. And gunicorn's default sync worker serializes streaming responses. Six concurrent streams against one worker produce a perfect staircase of first tokens, each arriving as the previous stream finished. The worker config was the queue; the models were waiting behind it. Switching to threaded workers (gthread, 16 threads) landed four of six first tokens inside a 1.4-second window.
One footnote on credentials: the docs insist a model access key and an API token are different, and they are, but a dop_v1_ API token authenticates against the inference endpoint anyway. Use the narrow key. A leaked model access key costs you some inference spend; a leaked API token costs you the account.
| Model | TTFT | Total | Cost per run | Empty runs |
|---|---|---|---|---|
| mistral-3-14B | 533 ms | 3.3 s | $0.000108 | 0/9 |
| llama-4-maverick | 676 ms | 17.9 s | $0.000362 | 0/9 |
| deepseek-3.2 | 869 ms | 5.4 s | $0.000416 | 0/9 |
| openai-gpt-oss-20b | 1792 ms | 4.6 s | $0.000235 | 2/9 |
| openai-gpt-oss-120b | 4797 ms | 16.7 s | $0.000367 | 0/9 |
| qwen3.5-397b-a17b | 9332 ms | 32.0 s | $0.000995 | 7/9 |
Medians across nine runs per model, max_tokens=512, one region, one evening. Time to first token ranged from 533 ms to 9.3 seconds. That 17x spread decides whether a feature feels instant or broken, and there's no way to guess it from a model card. Mistral 14B won on every axis: fastest to first token, fastest overall, cheapest per run, zero failures. The expensive models bought nothing for this workload. Total bill for all 54 calls: $0.0185, which is less than the time spent reading the pricing page.
The numbers hid a failure
Look at the table again and a different pattern shows up. llama-4-maverick is second fastest to first token at 676 ms, then takes 17.9 seconds to finish. openai-gpt-oss-120b is even worse: 4.8 s to start, 16.7 s total. TTFT and total time rank completely differently, and both matter depending on what you're building.
The last column in the table didn't make sense until someone looked at what actually came back. All 54 calls succeeded. No exceptions, no non-200s, no timeouts. Nine of them returned no readable text at all. qwen3.5-397b-a17b did it seven times out of nine. It's a reasoning model. It spent 487 of its 512 token budget thinking, ran out of room before writing a single word of the answer, and streamed all of that thinking into delta.reasoning_content, which is not part of the OpenAI schema and is invisible to every OpenAI-compatible client, this one included. The request succeeds. The tokens get billed. The dashboard says everything is fine.
completion_tokens: 512. reasoning_tokens: 487. content: 0 characters. You can pay full price for silence and have your monitoring call it a success. The general version of this failure is worse than the specific bug: if your evaluation watches latency and status codes, it is structurally incapable of seeing it. The check that survives a provider swap is on the payload itself. Zero characters of content against a non-zero billed completion is a failure whatever caused it.
The comments on the race post added the same point from three angles. When I test models on bursty side-project traffic, cold starts dominate p99 far more than tokens per second, so the winner is the model whose startup doesn't eat the latency budget, not the one with the best spec sheet. Same story with ranking: I've had a model place second on TTFT and finish fifth on total time because its decode speed was poor. And cost per call hides the quality floor: the cheapest model per request is rarely the cheapest per accepted result once repair prompts and retries enter the workflow.
Streaming a 125B MoE on a laptop
slotstream attacks a different constraint: model size on consumer hardware. Qwen3.8-Flash-Next is a 125B-parameter mixture-of-experts, 105 GB on disk at 4-bit. On a 48 GB Mac, the stock MLX loader pushed the machine into 48 GB of swap before the first token. The fix is a memory design, not a faster kernel. Keep the 3.8 GB dense trunk resident, stream experts from SSD through a fixed pool of cache slots shared by all 48 layers, and size that pool to what the machine actually has. It's one Swift binary, no Python, and it speaks the Ollama and OpenAI chat APIs.
| Your Mac | slotstream takes | Warm decode |
|---|---|---|
| 8 GB | 8.1 GB | ~3 tok/s, and it warns you it will page |
| 16 GB | 10 GB | ~4 tok/s |
| 24 GB | 16 GB | ~8 tok/s |
| 32 GB | 22 GB | ~9 tok/s |
| 48 GB and up | 33 GB | ~12 tok/s |
Only the 48 GB row is measured on real hardware; the smaller tiers are estimates from its fitted curve. At ~12 tok/s, decode is readable but noticeably slower than the 50+ tok/s you'd get from a cloud GPU. The trade is running on a laptop you already own, and the engine re-checks memory every 15 seconds, shrinking the cache when other apps want it back and growing it when pressure passes. Greedy decoding output is byte-identical across cache sizes, and that equivalence is a standing test, so resizing never changes answers.
Two details stand out from the measurements. Engine start takes about 2 seconds because only the 3.8 GB trunk loads before the first token. And speculative decode with the model's own draft head is right 86% of the time, buying 1.24x decode speed (10.3 to 12.8 tok/s) when the expert cache is warm. The slow axis is prefill, not decode: an 8,000-token prompt waits about a minute before the first token on a 48 GB Mac, because every token is read before anything is generated. Follow-up turns only prefill what's new, so time to first token stays flat as the chat grows, which is the right design for a conversation.
People on the thread running this model through llama.cpp put its floor at 64 GB when the experts fit in memory, and resident options like Rapid-MLX are several times faster when the whole model fits on a 128 GB Mac. slotstream's trade is explicit: slower decode in exchange for hardware you already have, with every number published alongside its measurement method, including the experiments that failed.
Common pitfalls
Benchmarking through a queue and blaming the models. Six concurrent streaming requests against one gunicorn sync worker produced first tokens at 1.3, 5.3, 7.3, 8.3, and 9.1 seconds, each arriving as the previous stream finished. That's serialization, not model latency. Check your worker class and thread count before you build a caching layer for a problem sitting in your Procfile.
Trusting the model catalog. The
/v1/modelsendpoint returned 72 models; the account could call six. No availability flag, no tier field, nothing. If you populate a model menu from that endpoint, assume most of what comes back is unreachable and ship fallbacks rather than a dead default.Monitoring status codes and latency, not content. Nine of 54 calls returned zero characters of text, and all 54 were HTTP successes. If your eval watches latency and status codes, it's structurally blind to billed silence. Check the payload: zero content with non-zero
completion_tokensis a failure, whatever field the reasoning count lands in.Hand-picking layers to protect in mixed precision. The three intuitive hypotheses, task circuits, computation sites, and weight statistics, all failed to predict which layers recover accuracy. Use causal intervention to find them, or skip the search entirely and spend the budget on global granularity, which won by 21-52 points in the study.
Running reasoning models with a tight
max_tokens. A 512 budget, 487 tokens spent thinking, zero characters of answer, full price. If your model streamsreasoning_content, count it, surface it, and raise the cap, or your dashboards will celebrate requests that delivered nothing.
One thing to remember
The pattern across all four results is the same. Cutting inference cost moves the failure somewhere your current tooling can't see it: quantization damage that intuition can't predict, KV caches the OS evicts while you switch apps, reasoning tokens billed into a field no OpenAI-compatible client reads, and prefill time that dominates when weights stream from SSD. Before you trust a model card, a benchmark thread, or a clean status code, build the measurement that looks at the payload, the layer, or the worker queue.
The bottom line
If you're quantizing a model and deciding where to spend extra precision bits, stop protecting layers by intuition