Skip to content

The Local Inference Squeeze: GPU Shortages, $3K Franken-Servers, and Ternary MoE on CPU

#local-inference #llama-cpp #ternary-quantization #gpu-shortage #mixture-of-experts

The state of local inference hardware ​

Local inference hardware is in a strange phase. Three things happened close together: NVIDIA started shipping the RTX PRO 5500, a workstation card with 84GB of GDDR7. A Reddit builder posted a $3,000 home server with 128GB of VRAM, assembled from four used GPUs and a cheap EPYC. And RTX 5090 stock nearly vanished while street prices climbed.

These aren't separate stories. They're three answers to the same question: how do you get enough VRAM to run serious models without renting a cloud instance?

Underneath the hardware is a software shift that changes the math. The llama.cpp PR adding Maple 20B-A1B, a ternary MoE model, runs at 88 tokens/s generation on a plain Apple M4 CPU. No GPU involved. 88 tokens/s is a real-time chat experience on hardware people already own. That changes what "local inference hardware" needs to mean at all.

The consumer GPU squeeze ​

The 5090 stock situation deserves its own paragraph, because it's the backdrop for every buying decision right now. The tracking threads went from photos of store shelves to photos of listings. Prices have been going crazy, and the general mood is that it's about to get worse. I watched the same cycle with the 4090: once the flagship evaporates, the mid-range follows within a quarter.

The practical effect: the 32GB class, which is the practical ceiling for consumer VRAM, is getting harder to buy at anything near MSRP. If your local inference plan was built around a single flagship consumer card, that plan is now on a timer.

Meanwhile, the alternatives have never been more visible. Workstation cards with 84GB. Used data-center GPUs at fire-sale prices. Each path trades off money, power, and patience differently.

RouteVRAMApprox. costBest for
Consumer flagship (RTX 5090)32 GBMSRP, street price spikingSingle-GPU desktop, 32B-class models at 4-bit
Workstation (RTX PRO 5500)84 GBWorkstation tier, below flagship70B-class at 4-bit, agent-heavy pipelines
Used server (4x V620)128 GB~$3,000 all-inMaximum VRAM per dollar, MoE at 4-bit, long context

The community read on this is consistent: people who bought 5090s early feel lucky, people who waited are watching prices tick up, and both groups are now eyeing the same used-GPU listings. I found myself doing the same math.

The workstation escalation: RTX PRO 5500 ​

NVIDIA's workstation line is heading in the opposite direction from the consumer line. The RTX PRO 5500 Blackwell packs 84GB of GDDR7, fifth-gen Tensor Cores, and second-gen Transformer Engine support, aimed at LLM inference, agents, and generative AI work at a desk.

84GB is 2.6x the VRAM of a 5090. That's the entire 70B dense class at 4-bit with room for a serious context window, or several 30B-class models resident at once. For multi-agent pipelines and multimodal workloads, that changes which models you can treat as "always loaded."

The interesting signal is positioning. The 5500 sits below the flagship workstation tier while carrying more VRAM than most rendering workloads will ever touch. NVIDIA is shaping workstation SKUs around inference demand, not just compute. The marketing copy says as much: professional-grade AI performance for "multimodal agentic" workloads, right from your desk.

Quick Take: while GPU stock evaporates at the consumer tier, the biggest local-inference wins right now are model formats that don't need a new GPU at all.

The DIY route: a $3K, 128GB VRAM server ​

One builder took the opposite approach: why buy one new workstation GPU when four used ones fit in a server chassis? The build is a Huanandzhi D12D board, an EPYC 7452 (64 cores, bought used for $170), 256GB of DDR4 RDIMM, and four Radeon PRO V620 GPUs at 32GB each. Total tab: about $3,000.

ComponentCost
4x Radeon PRO V620 (32GB each)$1,400
256GB DDR4 RDIMM 2666$610
Huanandzhi D12D motherboard$410
EPYC 7452$170
ASRock 1600W PSU$220
Case, fans, misc~$200
SSDalready owned
Total~$3,000

The builder tried a Lenovo P620 workstation first and returned it. I get it: I've run into the same wall of proprietary Lenovo tooling, interlocked BIOS quirks, and management layers that make a "standard workstation" feel like an appliance you don't own. The Shenzhen-motherboard route is less polished and more forgiving at the same time. Parts are interchangeable. You know exactly what you're running.

VRAM per dollar is the headline. All-in, that's about $23 per GB. Counting just the GPUs, it's $11 per GB. Nothing new comes close to that ratio.

Performance on current software is respectable. With Qwen3.8-next-flash at W4A16 (Autoround, 4-bit), the server does about 1,300 tokens/s prefill and 70 t/s on code generation, 60 t/s on prose, at 128K+ context, with MTP speculative decoding on a vLLM fork. Prefill at 1,300 t/s means a 6,000-token prompt is fully processed in under five seconds. First impressions were rough, the builder was disappointed with 27B-class speeds until Qwen3.8-next-flash changed the verdict.

The power draw is the hidden tax. Prefill runs 700-900W. Decode settles at 500-600W. At $0.15/kWh, eight hours of heavy prefill is about $1.08. That's manageable, but the machine is also a space heater, and you need a serious PSU and probably a dedicated circuit. The 1600W ASRock is not overkill.

Key numbers from the build:VRAM per dollar: ~$23/GB all-in, ~$11/GB for just the GPUs Prefill draw: 700-900W (~$1.08 for eight hours at $0.15/kWh) Decode draw: 500-600W Throughput (Qwen3.8-next-flash W4A16): 1,300 t/s prefill, 60-70 t/s decode at 128K context

The software side: ternary MoE in llama.cpp ​

The hardware race is one half of the story. The other half: models are getting dramatically cheaper to run.

PR #27000 adds Maple 20B-A1B to llama.cpp, DeepGrove's ternary MoE. The name is the spec: 20B total parameters, about 1B active per token. 24 layers. 256 experts, 8 active per token. Weights are stored as ternary values, -1, 0, or +1, packed in TQ1_0 and TQ2_0 formats.

The architecture borrows from recent lineages: Q/K RMS norms after projection (Gemma 4 style), a SwiGLU gate clamped at +7 (DeepSeek style), sliding-window attention over 512 tokens interleaved with global attention at a 3:1 ratio, and a partial rotary factor of 0.5. The KV cache is ISWA-style. The token embeddings and output head are the only two dense tensors in the model, which is why they're forced to F16.

The quantization math is what makes it run on CPU. At roughly 2.5 bits per weight, the 20B model sits in about 6-7GB. The official TQ2_0 GGUF runs on an Apple M4 at 216 t/s prefill and 88 t/s generation. Compare that with the $3,000 server doing 60-70 t/s decode on a much larger model, and the direction is clear: efficient formats are closing gaps that hardware was widening.

What ternary weights actually buy you ​

If you're building for edge or on-prem deployments, this matters more than any GPU launch. A 20B-class model at 88 t/s on CPU drops the entry point for "good enough" local inference from a multi-GPU server to whatever laptop you already own.

The PR thread also contains a useful quantization guide. TQ1_0 and TQ2_0 are both lossless packings of the same ternary weights: convert one to the other and the files are hash-identical. The head is where quality leaks:

ChoiceRecommendationMeasured cost
Ternary packingTQ1_0 or TQ2_0, identical weightsNone, bit-identical round-trip
HeadF16 if RAM allows, else Q4_KQ4_K head costs ~1.5 PPL @512, ~0.7 @2048
EmbeddingsF16, Q8_0, or Q6_K all fineNegligible difference
Eval contextUse 2048+, not 512Global attention layers need context

The context finding is the subtle one. Perplexity at 2048 tokens is much better than at 512 across the board, because three out of every four layers use sliding-window attention. They physically cannot see past 512 tokens. Starve the model and it gets dumber; feed it context and the global attention layers start capturing long-range patterns.

The debugging saga that nearly buried the PR ​

The Maple PR almost died on a perplexity discrepancy, and the thread is a masterclass in how quantization debugging should go.

Early numbers showed TQ2_0 scoring roughly 7 PPL lower than a q2_0 requantization. That gap smelled like a kernel bug. The obvious suspect was ARM, because x86 numbers looked consistent across CUDA and CPU. When I ran the same q2_0 file on x86, CUDA and CPU agreed within 0.03 PPL. On ARM, the gap reproduced. Classic story so far.

Except it wasn't a kernel bug. The two sides compared notes and discovered their wikitext-2 files had different hashes. The "same" eval set exists in multiple versions, and the file difference alone accounted for 5.77 PPL. On pinned, identical files, CPU and CUDA landed within 0.28 PPL of each other. The novice move would have been to "fix" a nonexistent ARM kernel bug and ship a silent regression.

What I found encouraging: the thread insisted on evidence. Someone proved their q2_0 to TQ1_0 round-trip was bit-identical, eliminating the weights. Another person ran a full test-split pass overnight on an M4. Someone else asked for the exact launch command to rule out the harness.

Then there was the other kind of moment. I posted a few PPL numbers early on and they were borderline unusable. A reviewer caught that I'd run only 20 chunks, about 10,240 tokens, and called it out. Fair. With a model that uses 512-token sliding windows, partial evals are worse than useless. Reproducibility matters more than speed when you're reviewing a quant format.

Common pitfalls ​

1. Sizing memory from total parameter count on MoE models. Maple is 20B but only about 1B parameters activate per token. You need RAM or VRAM for all 20B of weights, but per-token compute tracks the active 1B. People read "20B" and rule out small hardware. The TQ2_0 file loads in around 6-7GB and runs on a CPU.

2. Treating TQ1_0/TQ2_0 as the same thing as q2_0. TQ formats are native ternary packings; converting between them is lossless. q2_0 is a general 2-bit path that re-quantizes the weights, and it measured a real quality penalty in this PR. When you see a comparison between "ternary" and "2-bit" results, check which format produced each number.

3. Comparing perplexity across different eval files. The Maple thread burned days on a wikitext hash mismatch worth 5.77 PPL. Hash your corpora, pin the exact file, and only then trust cross-backend comparisons. A few points of PPL difference are meaningless if the files differ.

4. Skipping the head quantization decision. A Q4_K head costs about 1.5 PPL at 512-token context compared with F16. For short-context agent loops, that's a real quality hit. If you're memory-bound on an 8GB machine, take the Q4_K head. If you have RAM, spring for F16.

5. Ignoring the power budget on used-GPU builds. The 4x V620 server pulls 700-900W during prefill. You need a beefy PSU, a circuit that can take it, and a cooling plan. At $0.15/kWh the electricity is about a dollar a day per eight heavy hours, but the heat is the part nobody budgets for.

Sources ​

One thing to remember ​

Every source in this cluster points the same direction: the cost of local inference is falling, but not where most people are looking. GPU stock evaporates at the consumer tier, NVIDIA sells VRAM as a feature at the workstation tier, and builders mine the used market for cheap gigabytes. Meanwhile, the software stack is delivering the real step-change. A 20B model at 2.5 bits per weight, running at 88 t/s on a laptop CPU, makes the hardware discussion subordinate to model format efficiency for a whole class of workloads.

The bottom line ​

If you're building a local inference box on a budget and need maximum VRAM per dollar, the used-GPU server route wins: 128GB for ~$3,000 all-in, against ~$11 per GB for the GPUs alone. Just budget for 700-900W prefill draw and the patience that an EPYC-plus-retired-data-center-GPU build requires.

If you're serving a