Appearance
The readout gap is real
A model fails a reasoning task. Most evaluations stop there: the answer is wrong, so the model must not understand. A batch of new papers argues this conclusion is premature. The failure might live in the projection from internal state to output token, not in the reasoning itself.
The strongest evidence comes from "Wrong Prediction, Right Answer." The authors noticed that hidden-state probes on Qwen3.5 models decode correct answers even when the model's own sequence scoring completely collapses. The logic survives in the activations. What dies is the readout: format biases, option-order effects, and position preferences distort the final score so badly that the correct answer loses.
In my own eval runs, I've seen the same effect. Flip the order of multiple-choice options and sequence scores wobble, while an internal probe keeps voting for the same answer. The model has an opinion. The output layer just fails to express it.
Two parameters, 25 examples
The "Wrong Prediction, Right Answer" fix is almost comically small. Fit an additive correction to the output scores, a shift and a scale. That's two parameters. Fit it on as few as 25 unlabeled examples, and Qwen3.5 models recover 9 to 34 accuracy points across reasoning benchmarks. That's the size of jump many papers treat as a major release. The same correction transfers to OLMo-2-1B and Llama-3.1-8B without retraining.
Two things about this matter. First, the correction is target-label-free: no human annotation, no ground-truth labels, just a distributional adjustment. Second, the recovered decisions hold on hard instances that lexical overlap can't explain, and they beat count-preserving permutation baselines. The recovered answers aren't surface shortcuts. The model actually worked the problem out internally.
Key numbers
- 2 parameters: a shift-and-scale on output scores recovers 9 to 34 accuracy points
- 25 unlabeled examples: enough to fit the correction, no labels required
- 2M versus 55M: hidden-state verifier matches a text-only verifier at roughly 1/27th the parameters
- pass@32: highest among soft-thinking approaches on the tested models
Quick Take: The token stream is a lossy projection of what the model knows. This batch of papers treats that projection as a fixable engineering problem, not a law of nature.
Reasoning without tokens
"Sof Latent Thinking" takes the radical step of removing the head entirely. During normal reasoning, the LM head projects hidden states through a vocabulary-sized matrix at every step. That projection is expensive, and it forces every reasoning step into a discrete token. A lightbulb moment in latent space gets squashed into whatever word is closest.
The Soft Latent Thinking method replaces the LM head with a lightweight projector during reasoning. The model rolls out autoregressively in embedding space, where each step stays continuous. No tokenization, no vocabulary bottleneck. The authors tested it on DeepSeek-Qwen-1.5B and LLaMA-3.2-3B, and it improves pass@k at every k while reducing per-step compute during chain-of-thought. Pass@32, the widest net, beats all other soft-thinking approaches tested.
For inference cost, the impact is direct: the matmul drops from vocab-size by hidden-dim to head-dim by hidden-dim. On models with large vocabularies, that's a real slice of per-token latency. And because steps stay continuous, the model isn't forced to commit to a word it can't take back. Continuous reasoning is a compression win and a quality win at the same time.
A verifier that reads activations, not text
Test-time verification has a cost problem. When a model generates 16 candidate solutions, a text-based verifier re-reads all 16. That's 16 extra forward passes of a separate model on top of generation. The HSRM paper asks why you'd re-read text at all when the generator already computed useful representations during generation.
HSRM extracts hidden states from a frozen generator at reasoning-step boundaries and feeds them to a small Transformer encoder that ranks candidates. It's trained from self-generated trajectories with outcome labels, so it needs neither human process supervision nor a large pretrained verifier. The results: 2M parameters that match or beat a 55M-parameter text-only energy verifier in 15 of 16 generator-dataset settings.
The efficiency insight here is that verification can be a side effect of generation, not a second pipeline. You already paid for those activations. Reading them costs almost nothing.
Giving representations volume
The stochastic LayerNorm paper is a methodological sibling to the others. Its complaint: distances between point-valued representations don't track downstream function. Two nearby states can produce different behaviors. Two distant states can behave identically. If you're going to say "layer 14 is where this model knows X," you need a similarity measure grounded in what the model actually does.
Their move is to give representations volume. A light-touch modification to LayerNorm adds isotropic Gaussian noise at each residual-stream read, then renormalizes. Overlapping stochastic representations induce overlapping downstream distributions, which brings representational comparison under information-theoretic tools like the data-processing inequality. One learned allocation parameter per residual-stream read distributes a fixed rate budget across the stack. The result reads like transformer blocks operating with learned finite precision.
The payoff is tracing which counterfactual distinctions survive MLP blocks and which get exposed to specific attention heads. On ViT-S and GPT-2 small, the method recovers head-specific sensitivity to token distinctions that line up with known attention motifs. This is more of a tool than a result: if you're going to probe hidden states, you need similarity measures that mean something functionally.
Self-modeling is a readout problem too
The self-modeling paper is the most sobering result in this batch. The authors benchmark LLMs on verifiable behavioral questions: would a prompt edit change the final answer? Current models show nontrivial but limited skill and make systematic mistakes on simple counterfactuals about their own behavior. RL on synthetic self-modeling data improves aggregate skill across three open-source families, with some transfer to held-out tasks.
Then the catch. The improvements don't constitute introspection. A model trained on behavioral questions gets better at predicting its own behavior, but that doesn't mean it has privileged access to its internal decision process. It might just be learning robust statistical correlations about how models like itself respond to inputs.
That's the same readout gap, inverted. The wrong-prediction paper showed hidden states know more than output text says. The self-modeling paper shows behavioral outputs can improve without internal access deepening. Both point the same direction: surface behavior is a shaky proxy for mechanism.
Same bottleneck, five angles
| Paper | Maneuver | What it shows |
|---|---|---|
| Wrong Prediction, Right Answer | Linear correction on output scores, 25 unlabeled examples | Sequence scores are biased; hidden probes aren't |
| Soft Latent Thinking | Replace LM head with a lightweight projector during reasoning | Reasoning works in continuous embedding space |
| HSRM | 2M-param ranker over frozen generator's hidden states | Verification can reuse activations instead of text |
| Stochastic LayerNorm | Noisy LayerNorm turns point representations into distributions | Similarity measures need functional grounding |
| Self-Modeling | RL on synthetic behavioral questions | Behavior gains don't imply introspective access |
The thread: the LM head is a bottleneck that distorts, discretizes, and discards. Each paper pokes a different hole in that bottleneck, whether by correcting its biases, routing around it, reading before it, measuring what it obscures, or testing whether self-knowledge survives it.
Common pitfalls
Don't treat sequence scores as ability scores. If you're evaluating reasoning on tasks with fixed answer formats, the output distribution carries strong structural biases. Permute the options, vary the prompt wording, and fit a two-parameter calibration before concluding the model can't do the task. A model you thought was broken might just have a biased readout.
Don't discard activations when building verifiers. Standard test-time compute pipelines generated 8, 16, or 64 candidates and re-read all their text with a second model. HSRM shows a tiny encoder over the generator's own hidden states does the same job with a fraction of the parameters and none of the extra forward passes. That's free infrastructure if you're already sampling multiple candidates.
Don't treat representation distance as behavioral distance. Two hidden states close in Euclidean space can drive very different outputs, and distant states can behave identically. If you're comparing layers or counterfactual edits, measure functional distinguishability, not just geometric distance. The stochastic LayerNorm approach is one way; behavioral probes are another. Point distances alone tell you little.
Don't trust a model's self-reports about its own behavior. Models can answer "would this prompt change your answer?" from learned behavioral correlations. That's useful, but it's not introspection, and it breaks exactly where you need it most: on counterfactuals outside the training distribution. When a model tells you how it would behave, test the actual behavior. It's cheap.
One thing to remember: when you evaluate a model, you're testing two things at once, the reasoning and the readout that turns reasoning into answers. Most failures blamed on the first are actually the second.
The Bottom Line
If you're building a test-time compute pipeline, adopt hidden-state verification now. Reusing the generator's activations with a small ranker beats re-reading text with a separate model, and the 27x parameter reduction makes it viable where a second model isn't.
If you're running benchmark evaluations, calibrate the readout before trusting failure rates. A two-parameter correction on a handful of examples can recover 9 to 34 points, which means uncalibrated benchmark scores are silently conflating reasoning failures with expression failures.
One thing to watch: discrete tokens may stop being the default reasoning medium. Soft latent thinking moves chain-of-thought into continuous embedding space while cutting per-step compute and improving pass@k. Within six months, expect latent-space reasoning to show up as an inference option in serving stacks, alongside greedy decoding and beam search.