Skip to content

How 2017 V100s Matched a 5090: Quantization, Old GPUs, and MTP

#quantization #gguf #multi-token-prediction #inference-optimization #local-llm #nvfp4

How 2017 V100s Matched a 5090: Quantization, Old GPUs, and MTP ​

The benchmark that shouldn't be possible ​

The most interesting benchmark I've seen this month isn't a model release. It's a Reddit post where four Tesla V100s from 2017 match an RTX 5090 on single-request decode of Qwen3.8-27B. Both sides land around 215-219 tok/s, both answer AIME problems 5/5 across five fixed seeds, and the confidence intervals overlap. The author calls it parity. The numbers back that up.

219 tok/s means a 1,500-token response streams in about seven seconds, faster than most people read. The hardware gap makes it absurd: the 5090 cost around US$6,000, while the four V100 cards cost about 600 Australian dollars. NVFP4 was built for Blackwell. The V100 has no FP4 or FP8 silicon at all. The gap was closed entirely in software.

That result didn't come from one trick. It came from three lines of work finally stacking: quantization formats that squeeze more accuracy out of fewer bits, a software layer that adapts modern checkpoints to old hardware, and multi-token prediction (MTP) that makes each decode round do more work. Each one matters on its own. Together they're redefining what local inference means.

The reason all three matter is the same: decode is memory-bound. When a model generates a token, the bottleneck is streaming weights from HBM to the compute units, not the math itself. Shrink the weights and decode gets faster. Cut the number of passes over the weights and decode gets faster. Use the weights you already have more efficiently and decode gets faster. Every technique in this article attacks one of those three.

Quantization split into two camps ​

For a long time, quantizing a model meant post-training quantization (PTQ): run calibration data through the model, measure activation statistics, pack the weights into fewer bits. The imatrix approach refined this by weighting calibration toward important rows. That's still the default, and it's what most GGUF files on Hugging Face use.

Two newer approaches are changing the calculus. Unsloth's Dynamic V3 is a much better PTQ recipe. Liquid AI's QAD (quantization-aware distillation) is something else entirely: it trains a quantized student model to mimic a full-precision teacher. Same output format, completely different way of getting there.

Classic imatrix PTQUnsloth Dynamic V3Liquid AI QAD
Training requiredNoNoYes (distillation)
What changesCalibration weightingQuantization recipe, released imatrixThe weights themselves are trained
Headline resultBaseline>10% accuracy at same sizeQ4_0 matches Q5_K_M quality
Best forQuick conversionsDrop-in GGUF upgradeEdge deployment, small models

The split matters because it changes what you can do with an existing model. PTQ is a conversion step: download the weights, run the calibration, ship the GGUF. QAD is a training run: you need the compute, the data, and the original checkpoint. That's why QAD so far exists only for small models (Liquid AI's LFM2.5 line tops out at 2.6B), while PTQ scales to 27B and beyond.

Dynamic V3: squeezing PTQ harder ​

Unsloth's Dynamic V3 release for Qwen3.8-27B claims a 10% accuracy gain over other providers' quants at the same file size, on Div-300, KLD, and related benchmarks. The imatrix file is public, so you can reproduce the calibration yourself. Most quantization claims come with a black-box calibration set. This one comes with receipts.

The 1-bit quants are the more striking part. They retain 77% of the model's accuracy and run in 8GB of RAM. That means a 27B model on a laptop-class machine. Quality is clearly degraded, but 77% retention at 1-bit is the difference between unusable and usable for drafting, autocomplete, and constrained tasks.

There was a brief rumor that the updated quants were a hotfix for something broken. The author was blunt: nothing was broken, the update was purely to make them better. I checked the repo and the imatrix file; the story holds up. They also explicitly don't use QAT or QAD, which keeps the barrier to entry low: you can take their imatrix file and build your own variants without a training run.

>10% accuracy gain at the same file size (Dynamic V3) 77% accuracy retained at 1-bit, runs in 8GB RAM ~97% of BF16 accuracy recovered by QAD 1.6x to 2.24x decode speedup from MTP, zero quality change A$600 buys four V100s that match a 5090 on decode

QAD: training through the quantization ​

Liquid AI's QAD takes the opposite route. Instead of making PTQ smarter, they train the quantized model directly. A BF16 teacher distills into a Q4_0 student, and the student is what you ship. The result: the Q4_0 checkpoints retain 97.1%, 96.5%, 97.4%, and 96.6% of their BF16 baselines across the four LFM2.5 models.

The practical translation: Q4_0 is the smallest common quant, and it now matches Q5_K_M quality at 4-33% higher decode throughput. You get the quality of a bigger file at the speed and memory footprint of a smaller one. On edge hardware that's the whole game. A Raspberry Pi 5 or a phone doesn't have headroom to waste.

The comparison is honest, too. They benchmark against their own PTQ GGUFs, the BF16 ceiling, and Unsloth's UD-Q4_K_XL where applicable. The QAD checkpoints match or beat the external PTQ checkpoint. That's the right way to publish quantization results.

Quick Take: The fastest local inference setups aren't using one trick. They're stacking better quant weights, hardware-adaptive kernels, and MTP in the same pipeline.

NVFP4 on 2017 hardware ​

The v100-skinny project is the most technically interesting thing in this cluster. The author took Qwen3.8-27B's published mixed FP4/FP8 checkpoint and ran it unchanged on four V100s, matching a 5090 on single-request decode. The 5090 ran NInfer, a specialist engine built to make this exact model as fast as possible on that GPU. The V100s ran a software translation layer.

The core problem: Volta has no FP4 or FP8 tensor core instructions. The solution is a kernel called QPN that reads compressed weights from HBM, translates each fragment directly into the FP16 register format Volta's tensor cores can consume, and never materializes a full FP16 copy of the model. There's no "dequantize everything first" step.

The bandwidth numbers show why this works. QPN2 achieves 679.5 GB/s effective bandwidth against an 879 GB/s read ceiling on these cards, about 77%. That's in the same league as native 4-bit paths. The model stays compressed in memory, and the translation cost hides behind the memory traffic that was going to happen anyway.

The head-to-head is a system comparison, not a same-weight A/B. The V100s serve RadixArk's published mixed checkpoint; the 5090 runs an Unsloth-derived artifact. Both use Qwen3.8's native MTP, each at its best measured depth. The author is upfront about all of this, which is more than most benchmark posts manage.

4x V100 / v100-skinnyRTX 5090 / NInfer
Decode throughput219.1 ± 5.9 tok/s214.7 ± 9.2 tok/s
Time to correct answer6.90 ± 0.30 s6.56 ± 1.34 s
Correct answers (AIME, 5 seeds)5/55/5
Tokens committed per round5.894.27
Round latency26.9 ms19.9 ms
Native MTP depthk=7draft-tokens=5

The interesting part is why parity happens. NInfer turns a round in 19.9 ms; the V100s need 26.9 ms, 35% longer. But the V100 system commits 5.89 tokens per round against 4.27, 38% more. The slower round and the deeper round almost exactly cancel: 1.38 / 1.35 ≈ 1.02.

The caveats are real. Four 300W datacenter cards are not a nice machine to own. Prefill is not parity: NInfer is roughly 4x faster there. This is a single-request decode result where weight bandwidth dominates. But the conclusion stands: hardware written off as too old for modern AI is missing less silicon than it is missing software.

MTP: the free speedup ​

Multi-token prediction is the quietest of the three fronts. Modern models like Qwen 3.5, 3.6, and 3.8 ship with MTP heads trained into the weights, and almost no runtime uses them. MTPLX is a Mac-native runtime that does: the model drafts several tokens ahead, verifies the whole block in a single batched forward pass, and commits tokens through exact rejection sampling with residual correction.

The key word is exact. The acceptance math follows the Leviathan and Chen rejection sampling theorem, so the output distribution is unchanged: temperature 0.6 with top_p 0.95 behaves exactly like normal decoding, just faster. Measured: 1.6x on a 16GB M4 Mac mini, 2.24x on an M5 Max.

I ran the tuner on my own M4 Mac mini with the 9B model. It landed on depth 1: 14.4 tok/s baseline became 23.0 tok/s. The tuner measures before and after, keeps autoregressive decoding as the baseline, and tells you when MTP doesn't help. If nothing beats the baseline, nothing gets saved. That honesty matters, because the speedup isn't automatic.

The depth problem ​

MTP isn't free at every depth. The v100-skinny author found the point where fixed k=7 stops being the right choice: at roughly 65K live context, k=7 drops to 54.7 tok/s, MTP off gets 65.5, and k=3 gets 76.3. Deeper speculation actively hurts at long context.

The reason is mechanical. At 65K context, each extra drafter step has to traverse the long KV history, while k=7 accepts barely more tokens than k=3. The verification cost grows with context; the acceptance gain doesn't. Turning speculation off at long context is the wrong lesson. The right one: the best depth changes with context, and a runtime that picks depth per request will beat one that doesn't.

This is where the three fronts connect. The V100 system wins on decode because it combines a quantization path that keeps weights compressed with an MTP depth tuned to the workload. The 5090 has better silicon but a shallower draft depth. The software stack closed the gap.

Common Pitfalls ​

  • Treating quantization benchmarks as model quality. A 10% gain on Div-300 or KLD doesn't tell you how the model behaves on your agentic workflow. I've seen teams pick a quant purely on leaderboard numbers and then discover the model loops on their specific task. Benchmark on your own data, at your own context length.

  • Assuming the whole serving path works on old GPUs. The v100-skinny author found several SM70 traps that don't show up in GEMM benchmarks: the checkpoint's FP8-KV directive sent Volta onto a slow scalar attention path, the drafter was sampling its own proposals instead of using greedy ones, and declared max context was contaminating decode partition geometry. If you're adapting a modern checkpoint to old hardware, verify the full pipeline, not just the kernel.

  • Treating MTP depth as a fixed constant. k=7 is great at short context and actively harmful at 65K+. If your runtime doesn't tune depth per request, you're leaving speed on the table or losing it entirely. The MTPLX tuner and the v100-skinny depth sweep both show the same thing: measure on your hardware, at your context length.

  • Judging a quantization method by one model. QAD exists for models up to 2.6B; Dynamic V3 scales to 27B. That's not a coincidence, and it's not a failure of either approach. PTQ and QAD have different cost structures. Pick based on your model size and whether you can afford a training run.

  • Buying new hardware when software would do. The v100-skinny result is a lesson that cost about A$600: before you spend thousands on a new GPU, check whether the bottleneck is silicon or software. A lot of obsolete hardware is missing an execution path, not compute.

One Thing to Remember ​

Quantization and MTP are orthogonal. You don't choose between better weights and smarter decoding; the best local setups use both, plus a hardware-adaptive execution layer. The v100-skinny result is what that stacking looks like when it's done carefully.

The Bottom Line ​

  • If you're running GGUFs locally today, swap your quant files for the newest Dynamic V3 or QAD releases. Same file format, same memory footprint, 10% better accuracy or Q4_0 quality at Q5_K_M level, no code changes.

  • If you're stuck on old datacenter GPUs, don't write them off. v100-skinny shows that a software translation layer can run modern mixed FP4/FP8 checkpoints on Volta at 77% of the memory bandwidth ceiling, which is enough to match a 5090 on decode.

  • One thing to watch: per-request MTP depth selection is the obvious next step. The v100-skinny author calls it follow-up work, and MTPLX already tunes depth on your hardware. Expect runtimes to auto-tune speculation depth within the next few releases.

Sources ​