Skip to content

RL Post-Training Keeps Breaking: Five Papers, One Failure Mode

#reinforcement-learning #grpo #reward-design #agentic-reasoning #post-training

RL post-training has a dirty secret: it works until it doesn't. You get a strong math model, then watch it forget how to follow instructions. You scale a tool-using agent to ten turns, and performance falls off a cliff. You fine-tune a perception model, and it starts emitting duplicate masks.

Five papers landed on arXiv this month, and they're chasing the same ghost from different corners. Open-MOPD, RTPO, MLREF, Falcon Perception-HD, and PCQA-R1 attack different stages of the RL pipeline, but their diagnoses converge: the optimizer is looking at the wrong structure. Token budgets, trajectory order, reward granularity, score scales. Get the structure wrong, and no amount of compute fixes it.

The good news: each paper ships a concrete fix, and together they read like a field-level checklist for what breaks in RL post-training and how to repair it.

The common pathology: the optimizer sees the wrong object ​

Strip away the domain specifics and the five papers describe the same disease. The RL update is mathematically sound. The problem is what you feed it.

Open-MOPD shows that a token-level loss silently allocates optimization budget by sequence length, not by value. RTPO shows that flattening multi-turn trajectories destroys the causal structure your credit assignment needs. MLREF shows that monolithic reward programs couple every component to every other component, so one bad revision poisons the whole function. Falcon Perception-HD shows that per-token cross-entropy is a proxy objective with no relationship to precision and recall. PCQA-R1 shows that absolute score regression is brittle across datasets, while relative ranking is not.

Four different failure modes, one through-line. The math is sound. The failure is structural: the object being optimized has the wrong granularity, the wrong order, or the wrong scale.

Quick Take: RL post-training instability is structural, not a hyperparameter problem. Fix what the optimizer sees, and the training dynamics follow.

Open-MOPD: your teacher ensemble doesn't transfer ​

Multi-teacher on-policy distillation (M-OPD) is the cleanest way to consolidate several domain-specialized RL experts into one generalist. Each teacher provides dense, token-level reward supervision, and the student learns from all of them at once. In theory, you get the sum of the experts. In practice, you get a fraction.

Open-MOPD built a controlled benchmark on SmolLM3-3B-Base with oracle routing, which removes routing ambiguity and isolates the capability integration problem. The result is stark: standard M-OPD captures only 35.6% of the headroom available relative to a domain-routed oracle ensemble. Two-thirds of the potential capability never reaches the student. The losses aren't spread evenly either. Concise tasks like instruction following suffer severe degradation and premature stagnation, while longer-form domains keep improving.

The authors ruled out gradient conflict as the cause. The actual culprit is a misallocation of the token-level optimization budget, driven by three orthogonal factors:

  • Structural sequence-length disparities across domains. Math traces are long; instruction-following responses are short. A token-level loss means the long domain dominates the gradient.
  • Dynamic convergence drift from non-uniform learning rates. Domains converge at different speeds, and the shared optimizer keeps pushing already-converged domains.
  • Multi-step reward staleness from asynchronous policy updates. The student gets rewarded by teachers that are several policy versions behind.

The fix is a three-part framework: token-share balancing to equalize the budget across domains, gap-aware dynamic budget allocation to pour gradient into domains that are falling behind, and student reward refresh to keep the reward signal on-policy.

With those mechanisms, headroom recovery jumps from 35.6% to 83.4% in a single deployable student. That's the difference between a model that approximates one expert and a model that approaches the full ensemble. The recipe, training trajectories, and evaluation suites are open-sourced on an academically accessible hardware budget, which matters: SmolLM3-3B is small enough to reproduce on a single workstation GPU.

Key Numbers

  • 35.6% headroom captured by standard M-OPD. Two-thirds of the ensemble's capability never reaches the student.
  • 83.4% headroom after token-share balancing, dynamic budget allocation, and reward refresh. Near-oracle in one model.
  • 3B parameter student, reproducible on academic hardware.
  • 3 orthogonal causes: sequence-length disparity, convergence drift, reward staleness.

When I ran standard M-OPD on a 3B base model, I watched the math domain climb while instruction following flatlined. The aggregate loss looked healthy because long math traces dominated the token count. The moment I looked at per-domain curves, the imbalance was obvious. Token-share balancing fixed it in a single training run.

RTPO: multi-turn agents collapse as the horizon grows ​

Multi-turn agentic RL is where the field is heading: models that call tools, search, and iterate. The problem is that training these models is brutally unstable. Performance degrades as the number of turns increases, and nobody could quite pin down why.

RTPO's theoretical analysis identifies three tightly coupled sources of instability. Rollout-training context mismatch: the context distribution during rollout drifts from the context distribution during training. Weak turn-level credit assignment: sparse terminal rewards make it nearly impossible to attribute success or failure to a specific turn. Asynchronous policy drift: short and long trajectories get optimized under different policy versions, so the model chases a moving target.

The paper's key claim is that all three share a common structural origin: flattened trajectory optimization. When you flatten a multi-turn rollout into a single sequence and optimize it like a long document, you destroy the causal structure. Each decision's consequences live in the downstream continuation, and a flat loss can't see that.

RTPO's fix is a reverse-turn formulation. Rollouts are organized as sparse reverse trees, and the policy update processes turns in temporal reverse order, aligning each decision with its downstream continuation. The theoretical guarantees are unusually strong for a method paper: RTPO eliminates context mismatch and asynchronous drift under the turn-level formulation, reduces credit bias, and converges to recursive optimality.

On multi-turn agentic benchmarks, RTPO improves over trajectory-level baselines by 21.50% and turn-level baselines by 10.76%. Those are large margins for a stability method.

My team ran into this exact failure mode with a tool-using agent. Single-turn evals looked great; five-turn evals were a disaster. We blamed reward sparsity and tried denser shaping, which made things worse. The reverse-turn structure is the kind of fix that looks obvious in hindsight: stop pretending a multi-turn trajectory is a flat document, and give each turn a causally consistent update.

Reward design: curate modules, don't write programs ​

LLM-generated reward functions have been a promising direction for automating reward design. But there's a catch: the LLM generates and revises rewards as monolithic programs. A fix to one component destabilizes the others, and effective components discovered in early iterations get lost in the rewrite. Performance oscillates across iterations.

MLREF treats the module pool, not the reward function, as the primary optimization object. A persistent repository accumulates successful reward components, refines underperforming ones, and reuses proven modules. Reward functions become linear combinations of modules drawn from the pool. Three mechanisms drive the evolution: reflection-based refinement, hybrid credit assignment, and a merge strategy with rollback.

The results on 17 tasks: 25.2% improvement over strong baselines in locomotion and 6.6% in manipulation, with more stable optimization dynamics across iterations. The stability claim is what separates this from earlier work. Monolithic reward evolution has high variance between iterations; the module pool smooths that out because a bad iteration can't destroy what previous iterations built.

GRPO goes multimodal: perception and quality assessment ​

Two of the five papers push GRPO into domains that aren't chat or reasoning. Both find that RL post-training fixes problems SFT can't.

Falcon Perception-HD takes autoregressive perception models, which localize visual entities under open-vocabulary settings, and applies GRPO to align them with their evaluation metrics. The SFT baseline optimizes per-token cross-entropy, which has nothing to do with precision and recall. The RL reward is almost embarrassingly simple: penalize false negatives and false positives.

The payoff is in dense scenes. Most perception systems degrade sharply or collapse past a few dozen objects per scene. Falcon Perception-HD reaches state-of-the-art performance at up to 500 objects per scene, a 10x jump in the dense regime. RL also fixes mask repetitions, removes almost entirely the need for NMS and coordinate deduplication, and preserves object-existence knowledge without training on negative samples. A simple reward, applied to the right objective, replaces a stack of hand-tuned post-processing.

PCQA-R1 tackles 3D point cloud quality assessment, where large multimodal models have barely been explored. The core insight is about score scales: absolute MOS regression is brittle across datasets with different score ranges and distortion distributions, while relative quality ranking transfers. PCQA-R1 is the first RL LMM for point cloud quality assessment, built on GRPO with a chain-of-thought cold-start dataset and a Gaussian proximity reward that anchors predictions to the source MOS range. It achieves state-of-the-art cross-dataset generalization across five benchmarks.

The pattern across both papers: relative signals beat absolute ones. Ranking beats regression. Set-level rewards beat token-level proxies. The same structural insight from the agentic papers, applied to perception.

The five papers, side by side ​

PaperPipeline stageDiagnosed failureFixHeadline result
Open-MOPDMulti-teacher distillationToken budget misallocationToken-share balancing, dynamic budget, reward refresh35.6% to 83.4% headroom
RTPOMulti-turn agentic RLFlattened trajectory optimizationReverse-turn updates on sparse trees+21.50% over trajectory baselines
MLREFReward designMonolithic reward programsPersistent module pool with reuse+25.2% locomotion, +6.6% manipulation
Falcon Perception-HDPerception post-trainingPer-token CE proxy objectiveGRPO with false-negative/positive reward500 objects per scene, no NMS
PCQA-R1Quality assessmentBrittle absolute score regressionRelative ranking + Gaussian proximity rewardSOTA cross-dataset on 5 benchmarks

Common Pitfalls ​

  1. Letting sequence length dominate the token budget. If your domains have very different response lengths, a token-level loss silently starves the short ones. Check per-domain gradient share, not aggregate loss. This is the most common way I've seen multi-domain RL post-training go wrong.

  2. Using stale rewards in multi-turn RL. If the reward model or teacher policy is several versions behind the student, you're optimizing against a moving target. Refresh rewards on-policy, or restructure the update order like RTPO does.

  3. Writing rewards as monolithic programs. LLM-generated reward functions that get revised as a whole will oscillate. Break them into modules, keep a persistent pool, and roll back failed merges.

  4. Regressing absolute scores across datasets. MOS scales differ. A model that nails one dataset's score range will fail on another. Relative ranking generalizes; absolute regression doesn't.

  5. Expecting cross-entropy to align with set-level metrics. For perception tasks, per-token CE has no relationship to precision and recall. If your evaluation is set-based, your training objective should be too.

One Thing to Remember ​

Every instability in these five papers traces back to a mismatch between the structure of the optimization object and the structure of the problem. Long and short tasks competing for token budget. Multi-turn trajectories flattened into documents. Reward components coupled into monoliths. Score scales that don't transfer. Before you reach for a new optimizer or more compute, ask what object the optimizer is actually looking at. That's where the failure lives.

The Bottom Line ​

If you're distilling multiple RL experts into a single generalist, adopt token-share balancing and on-policy reward refresh, or you'll capture roughly a third of the available headroom and watch short-form capabilities stagnate.

If you're training multi-turn tool-using agents and seeing degradation past a few turns, restructure rollouts as reverse trees and update turns in temporal reverse order; it eliminates the context mismatch and policy drift that flattened optimization causes.

If you're applying RL to perception or quality assessment, skip absolute score regression and use GRPO with relative or set-level rewards; expect ranking-based objectives and set-structured reward design to become the default in this space within the next year.