Skip to content

Post-Training RL Is Maturing: GRPO Fixes, Staged Pipelines, and the Miles Stack

#reinforcement-learning #grpo #post-training #rlvr #reward-modeling

The default recipe has three failure modes ​

The default post-training recipe is SFT, then reinforcement learning with verifiable rewards (RLVR), with GRPO as the policy optimizer and a verifier or rubric supplying the reward. It looks simple because it is. It's also wrong in three specific places.

A cluster of recent papers shows the GRPO advantage estimator rewards guessing on bounded-answer tasks, a single per-trajectory scalar can't assign credit across long horizons, and fusing distillation with RL into one loss makes the two signals interfere. The fixes are concrete, and one is almost embarrassingly simple: change the order of your training stages.

Paper / toolFailure mode it targetsFixHeadline result
SIGNBALANCEGRPO gives guesses the same advantage as reasoningSign-preserving, per-class scaled advantageBeats GRPO on bounded-answer math and search agents
DRACOOne scalar per trajectory can't credit-assign long horizonsDynamic rubrics, closed-form per-step redistribution+15.9 on AppWorld over base, no verifier
OPD-then-RLFused distillation + RL signals interfereTwo stages, switch on OPD validationBeats pure OPD, pure RLVR, and joint baselines
TCSTest cases are scarce and not discriminativeTwo-stage adversarial test generationImproves pass@1 on TACO and LiveCodeBench
PreferenceEKFPreference reward learning is sample-hungryKalman filtering in a low-dimensional subspaceBetter sample efficiency and calibration on D4RL
MilesPost-training infra doesn't stay stable at scaleSGLang + Megatron-LM, async RL, low-precision recipesDay-0 support for frontier models

GRPO's advantage estimator rewards guessing ​

GRPO works by comparing rollouts within a group. Each rollout gets an advantage whose magnitude comes from within-group reward statistics, usually a normalization. For most math and code tasks that magnitude tracks something real: rollouts that reason their way to the correct answer score high, and the policy learns to reason more.

The overlooked case: a rollout that lands on the correct answer by guessing gets the same high magnitude. The paper calls this the spurious advantage, and it shows up in three situations. Bounded-answer tasks with a small candidate set, where guessing is cheap and the verifier can't tell the difference. Open-answer sets that host bounded sub-cases, like a numeric range where most of the probability mass sits on a few values. And search agents whose budget opens many paths to the same answer, so reaching it says little about the quality of the path.

If you've seen a GRPO run on multiple-choice reasoning slowly collapse into shortcutting, or a tool-use agent start favoring cheap actions because they happen to end at the same state, this is the mechanism.

SIGNBALANCE fixes the magnitude directly. The advantage keeps the verifier's sign, uses a global scale instead of a group-relative one, and restores zero-mean balance with a stop-gradient per-class rescaling. No compositional pieces that can combine into noise. Across math and search-agent benchmarks at different scales, it matches GRPO on open-answer math and improves on bounded-answer math and search agents, exactly where the spurious advantage lives.

Practical rule: before you blame the model for getting dumber, check whether your answer space is small enough that guessing pays. If it is, the estimator is lying to you.

One scalar can't credit-assign a fifty-step trajectory ​

RLVR works when a task has a programmatic checker. Most long-horizon agent domains don't. No compiler, no exact-match, just a long trajectory that succeeded or failed, with heavy ambiguity about which step mattered.

Multi-criteria rubrics are the standard substitute, and they inherit the same problem: scored once per trajectory, compressed into a single scalar. Across tens of steps, one number tells the policy almost nothing about which action caused the failure.

DRACO attacks the problem at the scoring layer. It generates rubrics dynamically during training, tuned to the policy's evolving capability rather than fixed in advance. Rubrics are still scored once per completed trajectory, but that judgment is redistributed over the steps responsible for the annotated rubrics, producing differentiated per-step advantages for GRPO. The redistribution is closed-form, so there's no trained attribution module to add flakiness.

The numbers on AppWorld: 15.9 points over the base model, and 5.3 points over GRPO trained with a sparse ground-truth reward, despite DRACO using no verifier at all. On out-of-domain Tau-Bench it gains 5.3 points over the base even without a frontier judge. Rubrics don't beat ground truth. A well-distributed rubric reward beats a sparse ground-truth reward, in exactly the settings where a denser reward isn't available.

Key numbers: DRACO gains 15.9 points over its base model on AppWorld and 5.3 over GRPO with sparse ground-truth reward, using no verifier. SIGNBALANCE matches GRPO on open-answer math and beats it on bounded-answer math and search agents. OPD-then-RL beats pure OPD, pure RLVR, and every joint baseline on logic and math benchmarks. TCS improves pass@1 on both TACO and LiveCodeBench.

Sequencing beats fusing ​

When you have two complementary signals, dense token-level supervision from a teacher and sparse reward from a verifier, the obvious move is to combine them in a single objective. Prior work did exactly that: weighted-additive combinations of distillation loss and RL loss, or teacher-modulated rescaling of the RL advantage. Both count as what the paper calls joint baselines. Both lose.

The two-stage scheme is what it sounds like: run on-policy distillation first, then RLVR. The paper's analysis of pass@k, learning dynamics, and parameter updates gives a consistent explanation. OPD expands the student's coverage of teacher-supported solutions. RL sharpens the policy within that support. When you fuse the two signals, expansion and sharpening happen at the same time and interfere, so you get the worst of both.

There's a practical recipe buried in the paper. The OPD validation score is the switch signal: run OPD until its validation plateaus, then switch to RL. And OPD is a better cold start for RL than SFT, which flips the assumed pipeline ordering. You don't do SFT then fused training. You do SFT, then OPD, then RL.

Quick Take: The most actionable result this month is to stop fusing distillation and RL into one loss. Run them in sequence, switch when OPD validation plateaus, and let RL sharpen what distillation already covered.

RL can write its own tests ​

Code RL lives and dies by test cases. They need to be sound, so they match the intended behavior, and discriminative, so they separate correct from incorrect solutions. Good tests are scarce, and that scarcity caps how much executable feedback the solver gets.

TCS turns test generation into adversarial RL. The generator produces tests as counterexamples, conditioned on the solver's current failure modes. Two stages, both trained from a rolling policy-aligned buffer. Stage one generates tests consistent with the reference solution. Stage two restricts the buffer to current failure modes and learns to generate counterexample tests. Across TACO and LiveCodeBench this improves both pass@1 and inference-time answer selection, where the generated tests vote on which candidate answer to trust. The learned generator also selects among other LLM outputs, which makes it useful beyond the training loop.

The lesson generalizes: test data isn't a static resource. It's an adversarial target that should track the model's current weaknesses, the same way a good reviewer writes tests against what the code is most likely to get wrong.

Preference learning finally has an uncertainty signal ​

RLHF from preferences works, but it eats samples. Active learning helps if you know where your uncertainty is, and for large reward models you usually don't. Full Bayesian posterior inference over a neural network is computationally prohibitive, so most systems fall back to point estimates and hope.

PreferenceEKF sidesteps the intractability. It frames active preference learning as sequential Bayesian filtering and runs an extended Kalman filter inside a low-dimensional subspace of the parameter space. As new preference queries arrive, the posterior updates, and sampling parameters from that posterior to compute acquisition functions becomes cheap enough for real use. On D4RL and V-D4RL it gets better sample efficiency, runtime, and calibration than other Bayesian deep learning approaches, and the learned reward models produce competitive offline RL policies.

For anyone building preference pipelines, this is the missing piece between the reward model and the query budget. You no longer have to choose between sampling efficiently and knowing what you don't know.

Production post-training catches up: Miles ​

None of this matters if the infrastructure can't hold a trillion-parameter run steady for days. Miles is the clearest sign that post-training RL has a production layer now. It pairs SGLang for rollout with Megatron-LM for training, with an FSDP2 backend if you'd rather train a HuggingFace model as-is. Fully async RL, LoRA and multi-LoRA adapters that load straight into SGLang, and token-in-token-out support so you don't pay the detokenize-retokenize round trip between rollout and training.

Across the runs I've been part of, three things stood out. On large MoE models, routing mismatch between rollout and trainer used to silently destabilize training; R3 replays the expert routing recorded during rollout and overlaps the compute and communication, so the mismatch is gone. When an SGLang engine dies, Miles recovers it in place. No restart, no pause, which is the difference between a one-hour hiccup and a dead experiment. And the MXFP8 and NVFP4 recipes are the first low-precision RL I've trusted; FP8 runs I tried elsewhere diverged, while these hold convergence and cut memory enough to train at frontier scale on fewer GPUs.

Day-0 support has become the scoreboard. DeepSeek-V4, Kimi-K3, GLM-5.2, Inkling, and Nemotron 3 Ultra landed the day their weights dropped, and the hardware list now covers NVIDIA's GB200/B200 lineup plus AMD's MI355X. Whether you use Miles or not, it sets the bar: post-training frameworks are now expected to handle frontier models on release day, with recipes for GRPO, PPO, REINFORCE++, and on-policy distillation under one roof.

Common pitfalls ​

What trips people up:

  1. Running vanilla GRPO on bounded-answer tasks. If your answer space is small, the spurious advantage inflates guess trajectories. Check reward statistics per answer class before blaming the model, and don't try to tune the problem away with KL penalties. Fix the estimator.
  2. Fusing OPD and RLVR into one loss. Weighted-additive and teacher-modulated advantage rescaling both underperform the sequential recipe. Run OPD, watch its validation curve, then switch.
  3. Scoring a 50-step trajectory with one scalar and expecting GRPO to distribute it. It can't. Either redistribute the judgment in closed form, like DRACO, or add dense signals. Otherwise the policy learns that every step is equally responsible.
  4. Treating generated tests as static gold data. Tests that only match the reference solution don't push pass@1 as far as counterexamples aimed at current failure modes. Generate adversarially and update as the solver changes.
  5. Jumping straight to FP8 or 4-bit RL without a numerical-stability check. Precision-induced divergence is real. I've seen naive FP8 runs drift within a few hundred steps; the MXFP8 recipes with proper scaling hold convergence, and on MoE models you also need the routing replay or the run will destabilize at scale.

One thing to remember ​

The through-line across this batch of work is that the remaining RLVR gains live in the details around the estimator and the reward signal, not in the loss function. Keep the advantage estimator honest, distribute credit properly, sequence your stages, and the same GRPO core keeps working. Skip those details and the recipe fails in ways that look like the model got dumber when the estimator was lying to you all along.

The bottom line ​

If you're post-training a reasoning model on math or logic with verifiable rewards, adopt OPD-then-RL: on-policy distillation first, switch when its validation plateaus, then RLVR. It beats every joint-fusion baseline and costs you nothing but a checkpoint.

If you're training a long-horizon agent without a programmatic checker, stop relying on a single per-trajectory scalar. Use dynamic rubric generation with per-step redistribution in the style of DRACO, or your GRPO signal will be noise across tens of steps.

If you're planning frontier-scale runs, the estimator and precision choices are the next battleground. SIGNBALANCE-style advantage fixes and MXFP8/NVFP4 recipes will be defaults within two release cycles, and frameworks like Miles already ship both.