Appearance
Post-Training Reasoning Models: Verifiers, Distillation, and a Bayesian Lens
The signal problem
Frontier reasoning models are made in post-training. The pretrained base supplies fluent language; everything that makes a model look like it can think comes later, from supervised warm-ups, RLVR runs, and distillation. Three papers landed in the same window, each attacking a different layer of that stack, and together they draw a coherent picture: post-training quality is a signal problem, not a scale problem.
Each paper takes a different angle. One builds a verifier-guided RLVR pipeline on a 3B model. One shows that ICL, SFT, and RL are the same Bayesian operation. One finds that 8 curated examples can match a 17K dataset in distillation. The table below shows how they line up.
| Verifier-guided RLVR | Bayesian unification | On-policy distillation | |
|---|---|---|---|
| What it studies | Reward construction and verification | The shared math under ICL, SFT, RL | Data selection in distillation |
| Core method | QLoRA + symbolic routing + RLVR | One Bayes posterior template for all training modes | 1-shot and 8-example training |
| Headline result | P3 up 21 points, P1 held flat | One framework that explains why these methods differ | 8 hard examples match a 17K dataset |
| Who should care | Anyone running RLVR on small models | Teams debugging unstable RL training | Anyone distilling into a smaller student |
Verifier-guided RLVR: rewards you can audit
The first paper constructs an explainable-reasoning pipeline on Qwen2.5-3B-Instruct. At 3B parameters, the model runs on a single consumer GPU, so the whole recipe is reproducible without a cluster. The pipeline has three stages: gold-anchored QLoRA supervision, task-aware routing, and group-relative RLVR.
First, supervised fine-tuning uses field-weighted QLoRA anchored to authoritative answers, which gives the model a solid starting point for each domain. Then a lightweight router sends logic problems to a first-order logic verifier built on Z3, and physics problems to a formula- and unit-aware symbolic solver. The router is symbolic, not learned, which keeps it cheap and inspectable. Verifier feedback then drives three things: candidate evaluation, self-revision, and reward construction during RLVR.
The reward design is the part to copy. Verifier feedback is not collapsed into a single scalar. It is decomposed into three scored dimensions: P1 for answer correctness, P2 for evidence or unit consistency, and P3 for reasoning depth and explainability. On 438 held-out examples, a small eval set so treat the gains as directional, RLVR lifts P3 from 50.68% to 72.20%. The model went from producing structured, checkable reasoning about half the time to doing it nearly three quarters of the time. Hybrid P1 stays roughly flat at 55.94%.
Key numbers
- P3 (reasoning depth): 50.68% → 72.20% after RLVR on a 3B model.
- P1 (answer correctness): ≈55.94% with hybrid verification across 438 held-out examples.
- Self-consistency alone: model-only P1 moves from 48.86% to 50.23%; the symbolic verifier adds the rest.
Read that split carefully. RLVR did not make the model more correct. It made the model's reasoning more structured and more legible. Reliability comes from the system layer at inference: gold-free self-consistency aggregates several candidates, then an optional question-only physics verifier applies conservative corrections. The design separates two problems that usually get tangled: the neural policy learns to write checkable reasoning, and the symbolic layer does the checking.
A Bayesian map of the whole post-training stack
The second paper is a theory note with a practical punchline. The authors put ICL, SFT, KL-regularized RLHF/RLVR, on-policy distillation, reward-weighted SFT, reward-weighted ICL, and advantage-weighted SFT on the same footing. The template is two steps. Build a generalized Bayes posterior over outputs given a context, using a reference model as prior and a utility signal (log-likelihood, reward, or advantage) as the evidence. Then approximate that posterior by forward-KL projection, either in-weights (SFT, RL) or in-context (ICL).
The equivalences hold at the level of objectives and first-order updates. They break on the source and granularity of the learning signal. That distinction matters more than the equivalences.
Once you see the template, some puzzling results stop being puzzling. Few-shot prompting sometimes hurts RL-tuned reasoning models. The template explains why: the few-shot prompt induces an in-context posterior approximation that can diverge from the posterior encoded in the trained weights. The two projections disagree, and the prompt steers you toward the wrong one.
The cold-start result matches what I've seen running RLVR on small models: reward runs on a base model produce long, empty chains that pad token count without adding reasoning steps. The Bayesian view names the mechanism. Importance-weighted KL projections require the reference policy to have mass under the good outputs, and a base model assigns near-zero probability to long chains of coherent thought. Supervised warm-up is not a training nicety. It is what makes the projection tractable in the first place.
Quick Take: all three papers say the same thing from different angles. Post-training quality is a signal problem. Fix the reward, the data, or the projection, and scale stops being the bottleneck.
The paper also reframes o1 and DeepSeek-R1 style systems as test-time Bayesian search combined with training-time KL amortization. The model learns to make the search's posterior cheap to approximate, then reruns that search at inference.
On-policy distillation: 8 examples are enough
The third paper studies on-policy distillation (OPD) from a data-centric angle. Train the student on exactly one example, and the run is consistently effective across every sampled training example. Harder examples give larger gains.
The paper then asks what actually drives the student's improvement. Token entropy is not the driver. Problems with high-entropy tokens do not produce better students. Harder problems generate longer chain-of-thought paths, and longer paths are what help. Long CoTs keep the student aligned with the teacher over a longer reasoning horizon and expose thinking patterns that short CoTs lack, such as self-correction. The paper points at the word "Alternatively" as a surface marker of reflection that only shows up in long CoTs.
Given that, the proposed selection rule is simple: keep only the hard examples. Even examples the teacher cannot solve still work, because the student learns the reasoning pattern before the answer goes off the rails. Across student models from 1.5B to 7B, all comfortable on consumer hardware, training on 8 selected hard examples matches the performance of the 17K dataset baseline. Eight examples. That is the difference between a day of data collection and a training run that finishes during lunch.
Where does this leave data collection? The OPD result echoes the verifier paper: the training signal matters more than its volume. A curated handful of examples that force long, self-correcting reasoning teaches more than tens of thousands of easy ones.
Where the three papers converge
Each paper improves a different component, but the pattern is identical. The verifier paper improves the reward by decomposing it. The distillation paper improves the data by selecting for reasoning length rather than entropy. The Bayesian paper explains why both work: every post-training method is a projection toward a posterior defined by some signal, and the signal's quality sets the ceiling.
The division of labor is the thing to internalize. The neural policy learns structure; the symbolic system enforces reliability. The chart below shows how little the neural side contributes to answer correctness compared to the system layer.
Look at the gap: sampling more candidates buys less than 1.5 points, while the symbolic check adds about 5.7. If you are building a reasoning pipeline, budget for both sides. No amount of RL turns a 3B model into a reliable calculator, and no verifier turns an unstructured explanation into a sound proof.
Common pitfalls
I've burned hours on each of these, and they map cleanly onto the three papers.
Skipping the SFT warm-up. The Bayesian analysis says cold-start is practically unavoidable for importance-weighted KL projections. A base model has near-zero mass on long coherent chains, so RLVR from scratch collapses or produces verbose garbage. Warm up first; it is not a formality.
Collapsing verifier feedback into one scalar. The verifier paper's P1/P2/P3 decomposition is what let RLVR improve explainability without sacrificing correctness. Merge everything into a single reward and the model games the correctness signal, ignoring reasoning structure.
Filtering distillation data by entropy instead of reasoning length. The OPD paper tested this directly: token entropy does not predict student gains, long CoT paths do. Curate for difficulty and for the reflection patterns you want copied, not for perplexity.
Assuming more data is better. Eight curated examples matched a 17K dataset. If your distillation run spends hours on a large pile of easy examples, you are paying for noise.
Expecting the neural policy to carry correctness alone. Self-consistency moved P1 by under 1.5 points; the symbolic verifier supplied the rest. Designing a symbolic safety net beats trying to RL the model into never making arithmetic errors.
One thing to remember
Post-training is not a battle of methods anymore. SFT, ICL, RLVR, and distillation are all doing the same Bayesian thing under the hood. The question that separates good runs from bad ones is always the same: what signal are you projecting toward, and how good is it?
The bottom line
- If you are training a small reasoning model, 3B or smaller, in a domain with externally verifiable answers, adopt verifier-guided RLVR with decomposed rewards, because you get the reasoning-structure gains of RL without gambling answer accuracy.
- If you are distilling a large teacher into a smaller student, stop collecting more data and start selecting hard, long-CoT examples, because 8 of them can match a 17K dataset baseline.
- If your RLVR runs keep stalling or diverging, audit the SFT warm-up before touching the reward, because the Bayesian analysis shows cold-start supervision is what makes importance-weighted KL projection tractable.