Skip to content

Local LLMs crossed the threshold: what changed this month

#local-llm #qwen-3-8 #inference #quantization #open-weights #hardware

The threshold crossed ​

Every few months, the local LLM crowd declares the threshold crossed. This time the evidence is harder to wave away.

I wired Qwen 3.8 27B into our Codex setup to compare against GPT Luna, our usual workhorse. Coding came out comparable. Then I ran it on one of our OCR pipelines and it beat Gemini 3.5 Flash Lite. That's a line item companies pay real money for, not a toy. Our team is now seriously discussing buying our own hardware, with estimates that it pays for itself in under two months.

That post landed the same week as two other data points that frame the moment. A solo developer shipped a 250M model that runs in 60 MB at 400 tok/s on a laptop CPU. Someone hosted Kimi K3, a 2.8T parameter model, on eight B300s for $190 per million output tokens. The spread between those three points is the whole story of local inference right now: tiny models that cost nothing, mid-size models that beat last year's frontier, and giant models whose economics finally make local hardware look sane.

Key numbers:

  • 400 tok/s: 250M model on a laptop CPU, no GPU, 60 MB total deployment. Faster than you can read.
  • 80 tool calls: one Qwen 3.8 27B prompt, zero human intervention, on a single RTX 3090.
  • 39,398: reasoning tokens Qwen 3.8 burns on one task at xhigh, vs 4,418 at low. That's why your context window evaporates.
  • $190: per million output tokens for Kimi K3 on 8x B300, 3.3x cheaper than the 1-bit GGUF alternative.

Qwen 3.8 27B: what a week of testing showed ​

The community verdict thread scanned roughly 2,000 posts across r/LocalLLaMA and r/LocalLLM. The consensus pick: a 27B dense multimodal model that moved the bar for local agentic coding. The strongest controlled evidence is tool-calling reliability, not a benchmark score.

One reporter ran a plain Python tool loop with no framework. Qwen 3.8 produced zero failed calls. Gemma 4 A4B and Qwen 3.6 A3B failed often. The same reporter rates 3.8 below both on raw code quality. Worse judgment, perfect plumbing. That's the profile you want in an agent.

Another user pulled a class schedule off a convoluted university website via 80 tool calls, zero human intervention. Single 3090, Q4 quant. A separate 1M+ token run built a full REST API and MCP server for a legacy forum from three prompts. It ran on a 16GB RTX 5060 Ti.

There are tradeoffs. Knowledge recall regressed versus 3.6, widely reported and best understood as a deliberate agent-first design choice. The model is trained to go search instead of recall. A legal-domain user got on-par-with-122B results once MCP and case access were attached. The knowledge didn't vanish; it moved into the toolbox. Vision works natively, including OCR-style reading of newspaper images at roughly 1,000 image tokens. But on a 16GB card at 64k context, it leaves as little as 150 MiB VRAM free. Keep text-agent and vision profiles separate, or offload the projector.

What the community is saying: the "IBM moment" framing keeps coming up. The argument goes that hyperscalers' moat was buying up hardware. Sanctions-driven competition has pushed small model quality up so fast that the mainframe analogy finally fits. I'm not sure the history is that clean, but the two-month hardware payback math is doing the heavy lifting, not the analogy.

The thinking dial changes everything ​

Most people get one thing wrong about Qwen 3.8: the shipped default is xhigh reasoning, and it thinks a lot. A lot. That's why your context window evaporates and your overnight job is still running in the morning.

Measured ladder on an RTX 5080 Laptop 16GB, llama.cpp, UD-IQ3_XXS quant, SVG generation task:

EffortReasoning tokensWall timeVisual score /25
Low4,418112 s21.8
Medium5,918127 s22.5
X-High39,398718 s24.0

That's 6.4x the wall time for +1.5 points on an eyeball task. But on pass/fail SWE-style tasks, xhigh went 9/12 versus 6-7/12 at lower efforts. The premium scales with whether the task has a verifiable failure.

The biggest controlled run of the week was a 67-hour, 40-arm benchmark across five production stacks. The headline: xhigh burned 7-11x more reasoning tokens than low for 0-4.7 extra points. In one head-to-head, low matched xhigh exactly at 89.3% on 1/7.5th the tokens. The author's verdict was harsh: "low is the rational preset. xhigh is for leaderboard screenshots."

I've run enough local models to know the temptation. You set it to max and walk away. But the data is consistent: medium for chat and analysis, xhigh only when there's a verifiable right answer. One user found xhigh took 7 minutes versus medium's 80 seconds on a bug-finding test, and caught every bug. Medium caught only the critical ones. Know which category your task falls in before you pick a preset.

Quick Take: Qwen 3.8 27B is the first local model where the reasoning preset matters more than the weights or the quant.

Quantization: the Q4 wars, mostly settled ​

The one controlled perplexity sweep on 16GB-fitting quants:

QuantSizePPLvs Q8
Q8_027.0 GB6.956100%
Q4_K_M17.1 GB6.95899.97%
IQ4_XS14.6 GB7.01399.2%
UD-Q3_K_XL12.5 GB7.11197.8%
NVFP4 (Q5K)14.4 GB7.20096.6%

Q4_K_M is effectively free on perplexity. But the community split hard below Q6 for complex reasoning. And you should take the "there are like 5 different Q4s and they are not equal" warning seriously. NVFP4 was the biggest disappointment: same size as IQ4_XS, worse perplexity.

The task-level benchmark tells a different story. At xhigh, every quant landed between 88.0% and 90.0% pass@1, with the 4-bit quants scoring above the FP8 baseline. Statistically a tie. The gap between quants is smaller than the gap between reasoning presets. The one failure that mattered: NVFP4 with reasoning off collapsed on HumanEval+ (13/30 vs FP8's 30/30). Never run reasoning off. It costs 8-12 points everywhere.

1-bit is comedy, not compute, for agents. Divergence hits 92% from BF16 by token 32. General chat survives; tool calls don't.

Hardware reality check ​

Reported speeds, attributed to the poster's hardware and runtime:

HardwareSetupContextSpeed
RTX PRO 6000 96 GBllama.cpp DFlash2, Q4_K_M262k153.9 t/s (304.9 with ngram on code)
RTX 5090 32 GBNVFP4-MTP-LOW262k121 t/s
2x RTX 3090vLLM + AutoRound INT4 + DFlash2131k120 narrative / 218 code
RTX 4090 24 GBUD-Q4_K_XL, MTP, Q4 KV262k40.7 t/s (59-60 with native MTP)
RTX 5060 Ti 16 GBUD-IQ4_XS + MTP-1, Q4_0 KV64k45.6 t/s
Strix Halo 128 GBQ8_0 + Q8 KV, ROCm, MTP142k9-19 t/s
Laptop CPUSHADOW-250M, <2-bit quant1M+400 t/s

Two patterns jump out. First, the "~200 tok/s" claims you see in headlines don't reproduce for you because they're measured at short contexts with Blackwell-tuned engines. Second, Windows costs you 10-15% on WDDM. I finally switched from llama.cpp on Windows to vLLM on Linux and got a 30-50% boost. That's the cheapest upgrade in this entire article.

If you're stuck at 16GB, you still have options. One user fine-tuned Gemma 4 12B specifically for tool calling and the CLI, because nothing else fits comfortably in their VRAM. They got a 2.7x improvement in tool-use success. Fine-tuning a smaller model for your specific workload beats squeezing a bigger one into the wrong shape.

At the other end, a user on an M1 Max said running the 27B at xhigh overnight for one task is impractical and no fun. They're waiting for the 35B A3B MoE, which will trade a little intelligence for a lot of speed.

The 60 MB counterpoint ​

While everyone argued about 27B quants, a solo developer trained a 250M model from scratch on 30B tokens of FineWeb. It's quantized under 2 bits. The whole deployment is 60 MB, smaller than a single high-res photo, and needs about 80 MB of RAM to run. It does 400 tok/s on a normal laptop CPU. No GPU.

The long context design is the interesting part. The most recent 2048 tokens stay in fp16 like a normal KV cache. Everything older gets compressed to 1 bit and written to disk, about 320 bytes per token. A million tokens of history is roughly 320 MB on disk. The model was trained from the start to retrieve from that disk cache, up to 100M tokens. It wasn't trained to reason over those tokens, only to retrieve and answer. In archive mode, it pulled a serial number that sat 50.6 million tokens deep in the archive.

The vocabulary is also unconventional: every token is a fixed 512-bit code, 8.4 MB for all 131k tokens, zero trained parameters. On WordSim-353 it scores 0.619 Spearman correlation versus 0.029 for random codes. The numbers are honest: cross entropy 3.15 nats per token, perplexity 23.3. It's a 250M model, it makes mistakes on open facts, and the author says so directly.

The architectural ideas are the point. Disk-backed context compression and fixed token codes look like a curiosity until someone ships them in a phone.

What big models actually cost ​

The Kimi K3 hosting write-up is the clearest look yet at frontier-model economics. Eight B300s on Modal at $56.79 per hour, tensor parallel 8, native MXFP4. Cold boot takes about 27 minutes (1.56 TB load, JIT, 51 CUDA graph captures). Steady decode at 92 tok/s, time to first token under a second. That works out to $190 per million output tokens. Left warm, it's $1,363 a day.

The alternative: Unsloth's 1-bit Dynamic GGUF at 594 GB, fitting on 8x A100-80GB via llama.cpp. At $19.99 per hour it's 2.8x cheaper on compute. But 9 tok/s and a 7-60 second time to first token push the cost to $620 per million tokens. 3.3x more expensive per token. Quality at 1-bit was fine: correct arithmetic, coherent prose. The economics are brutal though.

Put $190 per million output tokens next to a hardware purchase that pays for itself in two months. The hyperscaler moat starts looking thin. That's the math driving the "IBM moment" talk. At the extreme end, one homelab builder expanded to 36 DGX Sparks with 4.6 TB of unified memory. They run every notable model that lands, split into inference modules managed by a persistent agent. The API side is in flux too. An anonymous model called Ox Alpha appeared on OpenRouter with a 1M token context, free, and completed 8 of 10 DeepSWE tasks. Tokenizer forensics point at GLM-5.3 Flash.

Common pitfalls ​

Running xhigh by default. The shipped preset burns 7-11x the reasoning tokens of low for 0-4.7 points on most tasks. Set medium for chat and analysis. Reserve xhigh for tasks with a verifiable right answer.

Confusing budget with effort. llama.cpp's web-UI reasoning selector is a hard token cap that truncates mid-thought. It is not Qwen's native effort levels, which actually change how the model works. Use --reasoning-effort medium or the chat-template-kwargs form, not the budget slider.

Quantizing KV cache when you don't need to. f16 and q8_0 are not equivalent past 120k context, per one AMD tester. The working floor from the community: Q4 model weights with Q8 KV cache for agent loops. Don't go lower unless VRAM forces you.

Using 1-bit quants for agentic work. Divergence hits 92% from BF16 by token 32. Fine for chat, broken for tool calls. If you must, set presence penalty to 1.5.

Benchmarking on Windows. WDDM costs 10-15% versus Linux. The 30-50% jump from switching to Linux and vLLM is the cheapest speed upgrade in this entire article.

Judging quants by perplexity alone. PPL says NVFP4 is measurably worse than IQ4_XS. Task benchmarks say they tie. Both are true: perplexity measures token-level divergence, tasks measure whether errors get caught.

One thing to remember ​

The single most useful sentence from a week of community testing: the gap between quants is smaller than the gap between reasoning presets. Stop agonizing over Q4 versus Q6 and start thinking about what effort level your task actually needs.

The bottom line ​

If you're running agentic coding on a 16GB card, adopt Qwen 3.8 27B at Q4_K_M with medium effort and Q8 KV cache. The default xhigh will eat your context window and your overnight job.

If you're constrained to CPU-only or embedded hardware, the 250M class is the only game in town, and you should steal the disk-cache context trick for any long-context application.

If you're paying API