Appearance
Causal evidence is overtaking correlation
Mechanistic interpretability spent its first few years asking where features live. Attention heads that track entities. Directions in the residual stream that encode truthfulness. Probes that separate moral text from immoral text. All of that work produced a map, but nobody could agree on whether the map reflected the model's computation or the researcher's measurement choices.
Five papers from the August 2026 arXiv batch push the field past that argument. Each one builds an intervention: mask the head, patch the cache, reparameterize the coordinates, then watch what breaks. The consistent finding is that correlation-based attribution overstates what we know. What survives is a smaller, uglier, more honest set of causal claims about specific checkpoints.
The papers differ in what they target: attention heads, probe directions, cache segments, coordinate systems. The protocol underneath is identical. Localize. Intervene. Control. Confirm.
Key numbers
- 1.7-2.6%: share of attention heads that act as visual retrieval heads across 11 VLMs.
- 80 percentage points: grounding accuracy lost when masking the top 20 retrieval heads.
- 0.26 vs 0.013: mean pairwise cosine between moral probe directions and matched non-moral concepts.
- 72.5%: share of tabular experiments where removing low-importance heads first preserved performance.
- 95%: fraction of component-count decisions that flipped under a function-preserving reparameterization.
Retrieval heads were never text-only
The first paper asks whether vision-language models have an analog of retrieval heads, the sparse set of attention heads in LLMs that fetch stored knowledge and route it to output tokens. The answer is yes, and the numbers are stark. In eleven VLMs, about 1.7-2.6% of heads qualify, roughly 20 to 30 heads on a 7B-scale model. That's few enough to inspect by hand.
The authors call them Visual Retrieval Heads (VRHs). They found them by unifying existing head-scoring methods into one design space spanning query tokens, key aggregation, and cross-sample aggregation. Scoring attention from output prediction tokens against the sum over the ground-truth referent region won. Masking the top 20 VRHs on five referring-expression benchmarks drops grounding accuracy by up to 80 percentage points. The model that could point at "the red mug behind the laptop" now points at nothing useful. Masking 20 random heads barely moves the needle.
What makes this paper hit harder than a typical attention analysis is the follow-through. The heads generalize across visual reference tasks: they stay causal on attribute, spatial, counting, and visual-math benchmarks even though they were discovered through bounding-box prediction. They're functionally specific, corrupting localization while preserving output format. And they transfer across models: a VLM that shares its LLM backbone with another, but differs in vision encoder, projector, and instruction tuning, still relies on causally overlapping heads.
Practical reading: if you're debugging a VLM that mis-grounds text to images, you have a cheap diagnostic. Score heads from prediction tokens, mask the top 20, and see whether the failure reproduces. If it doesn't, the problem is upstream of retrieval.
Short version: across all five papers, the answers are narrower and more conditional than the original claims, and that's exactly why they're trustworthy.
Moral knowledge has a shared geometric core
The second paper probes how models organize moral knowledge. Six linear probes, one per Moral Foundations Theory category, trained on open-weight LMs. The question is whether the six resulting directions collapse into one moral detector, isolate from each other, or sit somewhere in between.
They sit somewhere in between, and the geometry is precise. The directions span near-maximal independent dimensions, so each foundation carries distinct information. At the same time they share a positive common component. Mean pairwise cosine of 0.26 across the six moral directions, versus 0.013 for a matched non-moral concept battery built identically. That gap is the integration signature: moral knowledge is organized as a shared hub with six spokes, not six unrelated axes and not one blob.
The structure is stable across architectures and scale, and it arrives early: the integration regime is reached well before probe accuracy saturates during pre-training. That timing matters. The model isn't learning six independent classifiers that slowly coordinate. It builds the shared component first, then refines the individual spokes.
The paper also pokes at Moral Foundations Theory itself. The authors find no evidence for the individualizing/binding split the theory predicts, though they admit the test is underpowered: only 20 candidate partitions exist to choose from. The geometry tracks corpus statistics instead. And on moral dilemmas, each dilemma's direction partially composes from its component foundations at 2.7 times a mismatched-pair baseline, while most of its variance encodes conflict-specific structure. The model represents moral tension, not a pre-resolved verdict. If you're doing safety alignment work, that distinction is the difference between steering a model's stated values and understanding the tension structure underneath them.
SCIT finds the cache carriers, checkpoint by checkpoint
Latent chain-of-thought models hide intermediate reasoning in continuous states instead of emitted text. Compact, but the causal object vanishes. You can't read the scratchpad, so you don't know what carries the computation.
SCIT, the Suffix Cache Interchange Test, is a direct answer. It builds exact source-recipient counterfactuals, patches declared segments of the KV cache, and asks which transformer object actually transfers the counterfactual computation. The protocol layers sufficiency tests with K/V component splits, hidden-state controls, semantic source controls, decoded validation, and matched corruption.
On CODI-GPT2 and a Sim-CoT-style GPT-2 reproduction, the finding is specific: counterfactual arithmetic transfers through value-cache suffix trajectories, not hidden states, not keys, not reusable answer slots, not single-token triggers. For the main CODI-GPT2 checkpoint, there's complete sufficiency-and-necessity evidence. The Sim-CoT-style checkpoint shows the same sufficiency pattern and passes decoded controls, but fails the matched-corruption requirement, so the authors won't call it necessary. That asymmetry is the good kind of rigor.
The bigger result is the carrier-regime map. Arithmetic-like cells in GPT-2 and 1B models preserve the latent-tail value/KV transfer. Competent 8B models and repaired non-arithmetic cells route through prompt-prefix or full-cache K/V instead. Boundary cells get no mechanism call at all. Where reasoning lives depends on model competence, not on a universal architectural law. If you're building on latent reasoning models, this is a direct warning: a mechanism identified in a small reproduction may simply not exist in the production-scale model.
Tabular attention heads break the vision-language pattern
Interpretability work clusters in vision and language. Tabular transformers get treated as a side quest. The fourth paper shows why that's a mistake: attention heads on tabular data behave differently from their image and text counterparts.
Using an importance-scoring metric across 40 diverse tabular datasets, the authors found that in 72.5% of experiments, removing the lowest-importance heads first caused the least performance damage, and removing the most important head first caused the greatest drop. That sanity check passes. But the fine-grained picture is unusual. Important heads are scattered across all six attention layers with no layer-specific trend. The bigger problem is that importance rankings don't transfer: across datasets with different schemas and feature spaces, the head importance profile changes substantially. A head that's critical for one table is noise for another.
That's the opposite of vision and NLP, where important heads cluster in identifiable layers and patterns recur across models. For tabular work, there is no universal pruning recipe. You have to measure per dataset. The upside is that the measurement is cheap and the code is public, so per-dataset scoring is a realistic step in any tabular transformer pipeline.
Representation measurements need basis-invariance checks
The fifth paper is the field's reality check on its own tools. Hidden coordinates aren't uniquely determined by a model's input-output function. Any function-preserving change of basis gives you a different hidden space with the same behavior. So representation-derived measurements should be invariant to those changes. Many aren't.
The paper shows that column-permutation parallel analysis violates this invariance. Its reference distribution and selected component count can change while the model function and the observed covariance spectrum stay fixed. The authors formalize why: a data-internal reference procedure can't simultaneously preserve every coordinate marginal, stay orthogonally equivariant, and remove cross-coordinate covariance. One of the three has to give.
The empirical toll: across five models, three retrieval domains, and 75 transformations, the median component-count disagreement is 0.79, meaning the number of "discovered" dimensions shifts by nearly one full component depending on hidden-coordinate choice. Median fixed-threshold decision disagreement is 0.26. In a centering-only control, 1,141 of 1,200 component counts changed despite an unchanged observed spectrum. Independent parallel analysis seeds changed none of the corresponding decisions. Orthogonally invariant comparator scores, by contrast, stayed numerically stable with similar held-out discrimination.
Translation: if a paper tells you "the model has k interpretable dimensions" based on parallel analysis, the k is partly an artifact of coordinate choice. Use orthogonally invariant comparators, or report the spectrum and skip the component count.
The five papers at a glance
| Paper | Object of analysis | Method | Headline finding | Practical consequence |
|---|---|---|---|---|
| Visual Retrieval Heads | Attention heads in 11 VLMs | Unified head scoring + causal masking | 1.7-2.6% of heads; masking top 20 cuts grounding by up to 80 points | VLM grounding failures can be localized to a handful of heads |
| Moral knowledge geometry | Probe directions in open-weight LMs | Six linear probes + geometry analysis | Shared component at 0.26 cosine vs 0.013 baseline | Moral knowledge is integrated; steering directions exist but encode tension, not verdicts |
| SCIT | KV cache in latent CoT reproductions | Suffix Cache Interchange Test | Arithmetic rides value-cache suffixes in small models; carriers shift at 8B | Mechanism claims must be re-verified per checkpoint and scale |
| Head importance, tabular | Attention heads across 6 layers, 40 datasets | Importance scoring + gradual removal | 72.5% of runs resist low-importance removal; no cross-dataset transfer | Measure per dataset; no universal tabular pruning recipe |
| Reparameterization invariance | Parallel analysis component counts | Invariance audit over 75 transformations | 0.79 median count disagreement; 1,141 of 1,200 counts flipped | Don't report component counts without invariance checks |
Common pitfalls
Four mistakes show up again and again, and these papers are concrete antidotes to each.
Treating probe directions as causal mechanisms. A linear probe that classifies moral text tells you the information is decodable, not that the model routes its decisions through it. The moral paper is careful about this framing; the VRH and SCIT papers both show that masking and patching reveal a different, sparser story than probing suggests. If you haven't intervened, you have correlation, not mechanism.
Masking heads without a matched baseline. The VRH result only means something because random-head masking was the control. In head-removal studies, performance drops as you remove any heads once you remove enough of them. Without a random or matched-corruption baseline, your "important head" finding is indistinguishable from "the model needs all its heads."
Assuming importance transfers across datasets or scales. Tabular head importance varies with schema. SCIT's carrier map varies with model competence. A ranking or mechanism found on one dataset, one model size, or one training run is a hypothesis about the next one, not a fact. Re-measure.
Claiming necessity from sufficiency alone. SCIT's Sim-CoT checkpoint passes sufficiency and decoded validation but fails matched corruption, so the authors withhold the necessity call. If your intervention changes the output, that proves sufficiency. It does not prove the component is required. Publish the sufficiency result as what it is.
If you remember nothing else
The map is not the mechanism. Probe directions, attention scores, and component counts describe the model's representation space, which is one of many function-preserving views of the same behavior. Causal tests with controls are what turn a measurement into a claim about how the model computes. Keep that distinction and most of the confusion in this field resolves itself.
What this means
If you're auditing a VLM for grounding failures, start by scoring attention from output prediction tokens and masking the top 20 retrieval heads. The VRH protocol is a few hours of work and will tell you whether the failure is in retrieval or upstream.
If you're building interpretability tooling, skip parallel-analysis component counts unless the method survives a function-preserving reparameterization check. Orthogonally invariant comparators stayed stable in the paper's 75-transformation audit, and they're the safer default for anything you ship.
One thing to watch: mechanism claims are about to get checkpoint-specific. SCIT's competence-gated carrier map already shows latent-tail reasoning in small models and prompt-prefix routing at 8B scale. Expect the field to stop accepting "we found the mechanism" papers that don't state which scale, checkpoint, and cache regime the claim holds for, and budget for re-verification when you scale up or change training recipes.