Skip to content

Local Inference Hit Its Tipping Point. The Bottlenecks Just Moved

#local-llm #llama-cpp #quantization #llm-inference #on-device-ai #mixture-of-experts

The year local inference stopped being a compromise ​

I rebuilt Ninfer from source and pointed it at Qwen 3.8 27B in NVFP4 on a single RTX 5090. Peak throughput was 220 tokens per second, averaging in the 170s over a long session. That's roughly double what llama.cpp gave me on the same card. A 4,000-token code file streams out in under twenty seconds, and the model never leaves the machine.

A week earlier, on two RTX 3060s, I pushed a 125B mixture-of-experts model to 400 tokens per second of prefill. Two 12 GB consumer cards from the previous generation. The fix was two settings I'd had wrong the whole time.

Local inference crossed a line this year. It's no longer the hobbyist compromise where you accept slow output to avoid an API bill. The interesting problems moved. They're now engine configuration, memory budgeting, and the question of who decides what gets installed on your device.

Key numbers

  • 220 tok/s on a 5090 with Ninfer and Qwen 3.8 27B NVFP4, roughly 2x llama.cpp throughput
  • 400 tok/s prefill on two RTX 3060s after fixing the split-mode flag (started at 36 tok/s)
  • 4 GB silently downloaded by Chrome as Gemini Nano weights on eligible machines
  • 0.49x modeled cost per token of a residential 8B fleet versus a data-center baseline

What actually fits on your machine ​

The useful way to think about local hardware isn't total GPU power. It's VRAM, because that's what quantization targets. The consensus, encoded in the CanIRun.ai hardware matcher and confirmed by my own tests, looks like this:

MemoryWhat usually fitsExample device
8 GB3B-9B chat models at Q4, plus tiny image modelsRTX 4060
12 GB12B-14B dense models, or smaller MoERTX 4070
16 GB20B-27B at Q4, comfortable 14B at higher qualityApple M4
24 GB30B dense models, mid-size MoE, local videoRTX 4090
32 GB+Larger MoE and high-quality image generationRTX 5090

An 8 GB card handles a 3B-9B chat model at Q4_K_M, the usual sweet spot between size and quality, plus a few tiny image models. A 7B-9B model at Q4 is roughly the size of a large game download and runs on a laptop GPU with room for a real context window. 16 GB opens up 20B-27B at Q4. 24 GB and up gets you 30B dense models, mid-size MoE, and local video generation.

The 32 GB+ tier is where the hardware economics get strange. The RTX 5090's launch price turned the running joke "5090 now officially costs $5,090" into a real decision point. I seriously evaluated an M5 Ultra Mac Studio with 256 GB of unified memory over buying a second 5090. 256 GB of unified memory holds a 100B-class MoE model plus a generous KV cache, all resident, no PCIe shuffling.

One trap the site calls out and that bites people constantly: MoE models load all experts into memory even when only a few are active per token. Active parameters tell you about compute. Total parameters tell you about memory. Confusing the two is how people buy the wrong hardware.

Quick Take: if your model doesn't fit at Q4 with room for your target context window, a bigger GPU is the expensive answer and a smaller active-parameter count is the right one.

The flag that was costing me 7x prefill ​

When I ran Qwen3.8-Flash-Next on my dual-3060 box, decode was fine at 14-15 tokens per second. That's reading speed. Prefill was stuck at 36 tokens per second, which meant an 8K prompt took nearly four minutes to ingest before the first output token appeared. I tweaked flags for days. Nothing moved.

So I benchmarked systematically. The answer was embarrassing in hindsight. llama.cpp on multi-GPU defaults to -sm tensor, which shards weights across devices in a way that silently disables op-offload, the code path that streams only the MoE experts a batch actually needs. With it disabled, every expert matmul fell back to the CPU. The instrumentation showed 1191% CPU usage and both GPUs at near zero.

Switching to -sm layer moved prefill from 36 to 135 tokens per second. Raising -ubatch from 512 to 2048 took it to 303. ik_llama.cpp doesn't have the split-mode cliff at all, and hit 407 at the same settings.

The two fixes multiply; neither works alone. -ubatch does nothing until the split mode is fixed, because the bottleneck was never the batch buffer. And the fancier flags I'd been toggling barely moved anything:

SettingResult
-rtr (ik runtime repack)159 t/s vs 407 t/s baseline
-ub 4096OOM at every context tested
--threads-batch 8 vs 16300 vs 303 t/s, no meaningful difference
KV cache q5_1 instead of q8_0freed too little VRAM to change anything

ik_llama.cpp's edge at -ub 2048 comes with a real cost. It pinned 106-108 GB of system RAM for a model that llama.cpp under --load-mode mmap ran from 21-32 GB. Same model, same machine. On my 128 GB box both fit. On a 96 GB box, that's the difference between running and not running, with almost no headroom for anything else.

Parallel slots are the same kind of trap. --parallel 8 dropped single-stream decode from about 12 to about 5 tokens per second, while aggregate throughput stayed roughly flat from 1 to 4 concurrent slots. Idle slots cost speed.

Small models got smarter without getting bigger ​

The most interesting thing I ran this spring wasn't a bigger model. It was a smaller one with a lookup table bolted on. Qwen3.8-Flash-Next uses Engram embeddings, which replace the static per-token embedding table with one indexed by the last 2-3 tokens. Hash the N-gram, fetch the vector, feed it into the network. O(1), constant time, no FLOPs. The model carries 51B parameters of N-gram embeddings while activating about 6B per token.

Why that matters: transformers burn early layers re-deriving static knowledge. "New York" gets re-assembled from subword pieces every time it appears. An Engram table memorizes those multi-token entities so the neural layers spend their depth on reasoning instead.

The community threads around this model had to kill a misconception, and I'll kill it here too. Engrams won't let you run a 1T model on a single server. The lookup is dumb. It keys on the last 2-3 tokens, your 200K token context has zero influence on what gets fetched, and the wider context can only accept or reject what the table returned. Push N from 3 to 4 and you dilute the training signal, because rare N-grams get less data. The table memorizes, the transformer reasons.

The main payoff is smaller models acting smarter. A 27B model used to split its parameter budget between reasoning and memorizing static patterns. Move the memorization into a cheap lookup table and every active parameter gets freed for reasoning. That's the architectural shift that matters, and it's why the local model list now looks like this:

ModelTotal paramsActive per tokenLocal reality
Qwen3.8-Flash-Next125B~6B93.7 GB GGUF; runs on dual 12 GB GPUs with enough system RAM
Mistral Small 4119B6.5Bsparse MoE, the local all-rounder pick
GLM-5.2753B40Bfrontier-class agentic coding, 1M context, serious hardware required
Qwen 3.8 MoE2.4T95B1M context; not a consumer-box model

What the community is saying: the LocalLLaMA threads this spring split into two camps. One wanted Engrams to be the magic that runs trillion-parameter models on a workstation. The other read the ablations and concluded the architectural gain is the story. The second camp has the evidence. The same threads have been full of appreciation for Unsloth's Daniel and Michael, who shipped GGUF support for the new architecture, including streaming the big N-gram structures from disk, on day one. That upstream work is why local inference moves at the pace it does.

The economics that decide where inference runs ​

Not all inference belongs on your box, and the question of where it belongs got a serious open stress test this year. HEARTH is a published feasibility study of running an inference fleet across millions of ordinary homes. The author modeled the idea five times, and the model changed its answer each time.

PassWhat changedVerdict
1Compared the facilities around the hardwareHomes win
2Added accelerator cost and utilizationHomes lose
3Used consumer and data-center hardware at quoted pricesHomes win
4Modeled production batching for a 32B modelHomes lose
5Tested different model sizesHomes win only for small models

The story the passes tell: facilities favor homes, but batching reverses the verdict. During generation a server repeatedly reads model weights while advancing many requests together. A larger batch spreads each weight read across more paying tokens, but every active request also eats KV cache. An 8B model leaves enough memory on a 32 GB consumer GPU for a useful production batch, so the consumer card's far lower purchase price matters. At 32B, the consumer card runs out of memory first, its batch stops growing, and the data center wins.

What survived is narrow: small-model inference where one node completes one request and no durable state is tied to one house. The model estimates the residential hardware-and-energy stack at about 0.49x the cost per token of the best data-center case, for a quantized 8B workload. That's a reproducible model output, not observed production, and the author says so repeatedly.

The idea can't be dismissed because US data centers used about 176 TWh in 2023, and Berkeley Lab's reference case projects 649 TWh by 2030. New campuses need substations, transformers, and a path to connect load the grid was never designed to carry. The US already has 82.5 million detached houses, each with a meter, an electrical service, and a building. Put modest compute behind existing connections instead of dragging a giant new connection to the compute.

The routing design is deliberately boring. A control plane assigns replicas of small models to qualifying nodes and distributes checkpoints outside the request path. The router knows which homes have the model resident, which are healthy, and which have batch capacity. During generation one home holds the KV cache and returns the tokens; nothing crosses the residential network mid-request.

The failure domain is the honest weakness, and it's the first thing I'd attack in a review. A data center's utilization is a smooth curve; a residential fleet is a step function. Homes drop mid-request, the in-flight KV cache dies with them, and recovery means replaying the prompt and paying prefill twice. For cheap retryable 8B requests that's tolerable. For anything that needs identical tokens on retry, it isn't. The Reddit discussion pushed exactly on this reroute tax, and the published model doesn't price it yet.

The sequencing lesson applies beyond this project: prove demand before deploying hardware. The revised plan starts with a pilot that contains no houses. Rent both hardware classes, benchmark the real serving stack, find a customer who'll pay for the workload. If the economics or the demand fail, stop before an electrician is dispatched.

The local-first argument has a clean version: your data stays on your device, so there's no privacy policy to write. That argument gets harder to make when vendors treat your disk as a delivery target.

This spring I verified something on a fresh Chrome profile that had received zero human input. No keyboard, no mouse, no Chrome UI, nothing. Chrome installed 4 GB of Gemini Nano weights into it anyway. The kernel filesystem log shows the whole sequence: directory created at 16:38:54, unpacker subprocesses spawning at 16:47:22, weights.bin moved into place at 16:53:22. Fourteen minutes and twenty-eight seconds, with no human action.

The details compound the problem. The directory is named OptGuideOnDeviceModel, internal jargon no user would map to "Gemini Nano LLM". The download is gated by the same rollout flag as the settings page where you could refuse it, so the install happens before the opt-out UI exists. Delete the file, and Chrome re-downloads it on the next eligible window. A deletion is a transient state to be corrected, not a directive to be respected.

That's the same playbook Anthropic ran with Claude Desktop, which registered a Native Messaging bridge in seven Chromium browsers without asking. Two vendors, same pattern: install a capability nobody requested, make removal harder than installation, re-install on every launch. At Chrome's scale the math gets ugly. 64% browser market share, 3.4 to 3.8 billion users, one 4 GB push per eligible machine. Between 6,000 and 60,000 tonnes of CO2-equivalent per rollout, paid by everyone.

The frustrating part is that on-device inference is the right answer for summarizing a page, extracting action items, or classifying a document. Apple's FoundationModels APIs show the version that respects users. The model is already on the device, the data is already on the device, and the API hands back typed values instead of Markdown you have to scrape:

swift
@Generable
struct ArticleIntel {
    @Guide(description: "One sentence. No hype.") var tldr: String
    @Guide(description: "3-7 bullets. Facts only.") var bullets: [String]
}

let response = try await session.respond(
    to: "Extract structured notes from the article.",
    generating: ArticleIntel.self
) { articleText }

Both approaches run models locally. Only one asked first. The difference is consent, and the local-first community needs to hold vendors to that standard. Local doesn't mean invisible.

Common pitfalls ​

Four things bit me this year, in order of how much time they cost.

Multi-GPU split modes. On llama.cpp, -sm tensor silently disables op-offload for MoE models, and every expert matmul falls back to the CPU. I lost a week to this. Use -sm layer. Getting it wrong costs about 7x prefill throughput.

Treating ubatch as a fixed constant. -ubatch 512, 1024, and 2048 gave me 135,