Appearance
Open any causal LLM and look at the attention weights on the first token. They're enormous, no matter what the token is. A BOS marker, a period, a random word, it hardly matters. Position one is a magnet. And if you look at the residual stream activations at that position, a handful of channels sit orders of magnitude above their neighbors. Those "massive activations" are a well-known headache for low-bit quantization.
The strange part is how long it took to explain why they appear. Four recent papers chip away at that question from different directions, and together they sketch a practical picture of model internals. One paper traces sinks to the causal mask itself. Another compares how transformers and Mamba carve up their representation spaces. A third builds task vectors using only forward passes. A fourth reads editable concept graphs out of vision transformers.
The thread connecting them: internal structure you can measure is internal structure you can change.
Attention sinks: not RoPE's fault
For a while, the standard explanation for attention sinks blamed positional encodings, RoPE in particular. The paper argues that's wrong.
The authors' experiments point to the causal mask as the driver. Because every token can only attend to itself and earlier tokens, attention self-concentrates at the start of the sequence. The first position attends to almost nothing but itself, and its value output barely gets mixed with anything. The result is a large-magnitude output at position one: the sink phenomenon, plus the massive activations that ride along with it. The mechanism shows up regardless of which token sits at the front, which is exactly what the sink behavior predicts.
That matters practically. If you're quantizing a causal model to int8 or lower, your calibration strategy has to account for structural outliers at early positions. Clip them away and you're fighting the architecture. Build per-channel or position-aware scaling and you get the range under control without losing precision everywhere else.
Mamba and transformers: same behavior, different geometry
The second paper asks a deceptively simple question. Mamba matches transformer quality on language modeling, but the internals look nothing alike. Does the representation geometry care?
It does, and the paper measures exactly how. Transformer residual streams are dominated by a single principal direction. Most of the variance concentrates in one axis. SSMs, by contrast, spread information evenly across dimensions. If you evaluate a hybrid architecture, the representation space gets visibly skewed after every attention layer, which suggests the attention block, not the MLP or the token mixer, creates the lopsided geometry.
The geometric divergence does not lead where you'd expect. Despite the contrast, effective capacity matches tightly between the architectures. Rank-constrained probes find that concepts occupy subspaces of similar dimensionality in both. And the dominant PCA direction in transformers does not carry more conceptual information than the uniform dimensions in Mamba. If your probe reads that direction as the concept, you're probably measuring geometry, not meaning.
When I first fit linear probes on transformer activations, I saw suspiciously high accuracy and assumed the model had learned clean concept directions. What I was actually seeing was the top PCA axis doing all the work. Whitening the residual stream before probing changes the results, and the honest version is usually less flattering.
The paper's closing observation is the one that stuck with me. At the level of local semantic manifolds, transformers and SSMs are highly aligned. Same topics, same nearest neighborhoods. Global divergence, local convergence. That's a license to transfer techniques like probing and steering between architecture families, as long as you control for the geometry differences first.
Quick take: Transformers and SSMs look nothing alike at the global geometry level and behave almost identically at the local semantic level, and that gap is where a lot of interpretability work actually lands.
Task vectors without the fine-tuning bill
Task vectors are one of the cleanest model-editing ideas around. Fine-tune a model for a behavior, subtract the pretrained weights, and the difference is a direction in weight space that adds the behavior back when combined with the base model. The catch is the fine-tuning. Every new behavior costs a training run.
Training-Free Task Vectors (TFTVs) remove that cost. The method maps activation steering vectors to rank-one weight-space edits using only forward-pass statistics. No gradients, no fine-tuning. The resulting directions satisfy the arithmetic properties you actually want: add to learn a behavior, subtract to forget it, compose multiple edits without retraining.
The evaluation looks at behavioral control on LLMs, things like sycophancy, honesty, and other traits you'd want to amplify or suppress. TFTVs consistently strengthen or weaken target behaviors while preserving general knowledge and problem-solving ability. Against other editing and steering baselines, they show stronger trait control with equal or better utility preservation.
I've spent GPU-hours fine-tuning checkpoints just to extract a task vector, then watched it fail to transfer because the fine-tune overfit. A method that replaces that whole loop with a few forward passes changes how often you can iterate. You test a behavior edit in minutes, and if it overshoots, you subtract it back.
Key numbers from this cluster:
- 11.0%: accuracy gain on Waterbird from concept-circuit intervention over prior spurious-correlation methods
- 0: fine-tuning steps needed for TFTVs, just forward-pass statistics
- 40% vs. 9%: illustrative share of residual-stream variance held by the top PCA direction in transformers vs. SSMs
- 2: complementary circuit views from Cross-Layer Transcoders, global (input-invariant) and instance (prediction-specific)
Concept circuits: reading a ViT's world knowledge
The fourth paper moves to vision. Vision transformers generalize impressively across domains, but nobody had a clean answer for where "world knowledge" lives in their weights. The paper's answer: it lives in concept circuits, and you can read them out with Cross-Layer Transcoders.
A concept circuit is a directed graph. Nodes are sparse, interpretable concepts, edges are interactions across layers. The method produces two views. The global circuit is input-invariant, recovered directly from the cross-layer weights, and captures the reusable knowledge baked into the model. The instance circuit is input-dependent and shows which concepts and pathways actually fired for a specific prediction.
The practical payoff is spurious correlation removal. On the Waterbird dataset, intervening on the instance circuit beats existing methods by 11.0% at steering the model toward correct predictions. The same circuits expose shortcut dependencies automatically, without manual inspection of failure cases. And comparing global circuits of CLIP vs DINO shows how different training regimes shape the structural wiring of the model.
That last bit feels like the start of a new comparison culture. Instead of comparing models by accuracy alone, you compare their circuits. You ask which concepts each model treats as load-bearing and where the shortcuts live.
Four papers, four control surfaces
Taken together, the papers cover four different places you can grab a model by the internals.
| Paper | What it studies | Method | Practical payoff |
|---|---|---|---|
| It's Not RoPE that Creates Sinks | Attention sinks and massive activations | Causal analysis of self-concentration and value-non-mixing | Outlier-aware quantization, position-aware calibration |
| Global Divergence, Local Convergence | Representation geometry, transformers vs SSMs | PCA across scales + rank-constrained probes | Safer transfer of probes and steering across architectures |
| Training-Free Task Vectors | Behavioral control via weight directions | Forward-pass statistics to rank-one edits | Model editing without fine-tuning |
| World Knowledge in the Weights | Concept circuits in ViTs | Cross-Layer Transcoders | Spurious-correlation discovery and removal (+11.0% Waterbird) |
Each one is a diagnostic that turns into a lever. The sink paper tells quantizers where the outliers come from. The geometry paper tells probers how to avoid fooling themselves. TFTVs make behavior edits cheap enough to iterate. Concept circuits make vision model failures explainable and fixable.
Common pitfalls
Quantizing without a fat calibration set. When I first calibrated int8 ranges on a 7B model, the early-token massive activations dominated the scale factor and everything else compressed to near zero. The paper explains why: the sink is structural, produced by the causal mask, so it shows up in any prompt. Calibrate on prompts long enough to expose early positions, or handle the affected channels separately.
Reading the dominant PCA direction as meaning. A transformer's top residual-stream direction holds a huge share of variance and almost no concept information. Raw activations make linear probes look flattering. Fit the same probe on whitened activations, and if accuracy collapses, you were measuring geometry. Report both numbers.
Applying task-vector edits at every layer. TFTVs are rank-one edits tied to the layers whose activation statistics produced the direction. Spread the edit across all layers and the behavior overshoots while general knowledge degrades. Apply it where you measured it, verify with the addition and subtraction arithmetic, and sanity-check utility on a held-out reasoning set.
Cutting shared concept nodes in instance circuits. The biggest risk when removing a spuriously correlated concept is that the same node carries legitimate signal for other predictions. Check the global circuit first. If the concept is reused across classes, you need a more surgical intervention, which is exactly the case where the Waterbird gains come from.
One thing to remember
Every finding in this cluster starts as a reproducible measurement on real activations, and you can rerun all of it on your own model. Your model has its own sink behavior, its own dominant directions, its own concept circuits. The papers provide the tooling, but the diagnostics only mean something when you apply them to the specific model you're shipping. Sinks are structural, so quantization must budget for them. Geometry differs across architectures, so probes must control for variance direction. Task vectors no longer require fine-tuning. And concept circuits turn spurious correlations into editable graphs.
The bottom line
- If you're quantizing or compressing a causal LLM, plan for massive activations at early positions before you calibrate. They're caused by attention self-concentration and value-non-mixing, not positional encoding, so outlier-aware scaling beats clipping.
- If you're iterating on behavioral control and fine-tuning is too slow or too expensive, switch to TFTVs. A few forward passes per behavior gets you additions, subtractions, and compositions with stronger trait control than steering and utility preservation that matches fine-tuned task vectors.
- If you're debugging vision model failures, adopt concept circuits now. The 11.0% Waterbird gain over prior methods is the strongest published result on that benchmark. Watch the CLT training cost: expect extraction of global circuits to get cheaper within a few months as the technique spreads.