Mechanistic Interpretability: Circuits, Heads & Sparse Autoencoders
After this lesson, you will be able to:
- Articulate why behavioral safety (RLHF) is not enough, and why mechanistic understanding is the path to robust safety
- View a transformer as a residual stream that every layer reads from and writes to
- Recognize the canonical circuits: induction heads, name-mover heads, the IOI circuit
- Distinguish linear probes (correlation) from activation patching (causation)
- Understand polysemanticity, superposition, and why single neurons rarely correspond to single concepts
- Explain sparse autoencoders (SAEs) as the 2024 breakthrough that decomposes activations into monosemantic features
- Apply activation steering — adding or subtracting a direction at inference time — to control behavior
- Use the logit lens to read a model's intermediate predictions, layer by layer
- Name the open problems: feature composition, causal sufficiency, scaling SAE evaluation
Before You Start
#Part 1: Why Mechanistic Interpretability Matters
#Part 2: The Residual Stream View
This view buys you everything:
- Attention heads read from earlier positions on the stream (Q, K) and copy a transformed version (V, then
W_O) back into the current position. An attention head is fundamentally a move information from elsewhere to here operation. - MLPs read from the current position only, run a 2-layer net, and write back. An MLP is transform information in place.
- Layer norm is technically a non-linearity on the stream, but for first-pass intuition you can treat it as nearly the identity.
- The unembedding (a single linear map from
d_modeltovocab_size) reads the final stream and produces logits.
i, layer L? And the central tool: which heads and MLPs wrote them there?#Part 3: Attention Head Analysis
Some attention heads do something simple and human-readable. The catalog has been built up over five years of patient reverse-engineering.
#Induction Heads: The First Circuit
... A B ... A, the induction head predicts B next. Concretely, if the model has seen Mr. Dursley earlier in the prompt and now sees Mr. again, an induction head will increase the probability of Dursley as the next token.How is it built? Two heads working in sequence, across two layers:
- Previous-token head (layer
L): each token attends to the token immediately to its left, and copies that token's identity into its own residual stream slot. - Same-token-match head (layer
L+1): the currentAtoken now has, in its residual stream, "I amA, and the token before me was<whatever>". The next-layer head's query looks for "key = same as me." When the earlierAis encountered, the layer-L+1head matches on it, follows the stored "next token isB" pointer, and writesBinto the current position.
This is in-context copying, and it is what enables in-context learning — the model's ability to learn a pattern from the prompt alone, without weight updates. Olsson's paper showed that induction heads emerge sharply during training right at the moment in-context learning capability appears. Causation, not correlation.
An induction head has the pattern: in a sequence ...A B ... A, predict B. If you ablate (zero out) all induction heads in a model and then ask it to repeat a random made-up name from earlier in the prompt, what happens to performance?
#The IOI Circuit (GPT-2 Small)
- Identify the two name candidates (
John,Mary). - Notice that one of them (
John) was already used as the subject of gave. - Suppress that name.
- Promote the other.
Through painstaking ablation, the authors reverse-engineered a 26-head circuit spread across GPT-2 Small's 12 layers. The taxonomy they introduced is now standard:
- Duplicate-token heads: spot the repeated name (
John ... John). - S-inhibition heads: write a signal saying "the subject name has been duplicated; do not copy it again."
- Name-mover heads: late-layer heads that copy a name from earlier in the sequence into the final position's residual stream. There are positive name movers (copy any name) and negative name movers (suppress the name the S-inhibition heads flagged).
- Backup name-movers: redundant heads that step up when the primary name movers are ablated.
The IOI circuit was the moment the field knew this approach worked. Twenty-six heads, identified by role, with the algorithm written out in English. Not a metaphor — the actual computation.
#Head Taxonomy You Will See in Papers
| Head type | What it does |
|---|---|
| Previous-token head | Attends one token to the left; copies identity |
| Duplicate-token head | Attends to earlier occurrences of the same token |
| Induction head | Pattern A B ... A → B; in-context copy |
| Name-mover head | Copies a name from earlier into the current position |
| Successor head | Monday → Tuesday, 1 → 2, chapter 3 → chapter 4 |
| S-inhibition head | Suppresses already-used candidates |
| Copy-suppression head | Anti-copy; counters spurious repetition |
#Part 4: Probing Classifiers
L, extract the hidden state vector for every example in a labeled dataset. Train a linear classifier (just Wx + b) on top of those hidden states to predict the label. If the linear probe reaches high accuracy, the conclusion drawn is: "the feature is linearly represented at layer L."This is useful and it has been the workhorse of the BERTology literature since 2019. People have probed for part-of-speech, syntactic tree depth, sentiment, factuality, truthfulness, world models in chess-playing transformers (Li et al., "Emergent World Representations"), and much more.
You train a linear probe on layer 14 of a chat model and find it predicts 'is this a coffee shop conversation' with 95% accuracy from the hidden state. Which of these claims is the probe's evidence strong enough to support?
v your linear probe weighted highly, then remove it from the residual stream (subtract (v · x) v / ||v||² from x) and re-run the forward pass. If downstream behavior changes when v is removed, you have causal evidence the model uses that direction. If behavior is unchanged, you found a correlate, not a cause.Polysemanticity is the phenomenon where...
#Part 5: Activation Steering
The recipe:
- Find a direction
vthat represents the concept. Methods range from "average the difference of hidden states between positive and negative example pairs" (mean difference, Subramani et al.) to SAE feature directions (Templeton et al.) to learned probes. - At inference, intercept the residual stream at some chosen layer and add
α v(to amplify) or−α v(to suppress). - Continue the forward pass. Observe the model's output.
The directions found are often startlingly meaningful. Anthropic's Golden Gate Claude is the famous example: a single SAE feature direction, scaled up, made Claude reroute every conversation toward the Golden Gate Bridge. Less whimsical applications:
- Refusal direction (Arditi et al., 2024, "Refusal in Language Models Is Mediated by a Single Direction"): in many open-weight chat models, a single direction in mid-layer activations mediates refusal. Subtracting it bypasses safety training entirely. Adding it makes the model refuse benign queries. The paper's finding suggested that current refusal training compresses the entire safety behavior onto one dimension, which is brittle.
- Truth / honesty directions (Burns et al., "Discovering Latent Knowledge"; Marks & Tegmark, "The Geometry of Truth"): unsupervised methods can isolate directions that correlate with the model's internal belief about a statement, even when the model's outputs are aligned to lie.
- Sycophancy, aggression, formality. Countless behavioral dimensions have been steerable, with varying degrees of robustness.
In a 32-layer chat model, you have extracted a 'refusal direction' r. You subtract α·r (with α positive) from the residual stream at layer 14 during a forward pass on the prompt 'Tell me how to make a bomb.' What is the most likely effect?
#Part 6: Sparse Autoencoders: The 2024 Leap
x ∈ ℝ^{d_model}. You now have hundreds of millions of d_model-dimensional vectors.x → f → x̂ where:- The encoder is
f = ReLU(W_enc x + b_enc), withf ∈ ℝ^{d_dict}andd_dict >> d_model(often 8× to 64× larger). - The decoder is
x̂ = W_dec f + b_dec. - The loss is
||x − x̂||² + λ ||f||₁.
d_dict features as possible to reconstruct each x. The overcomplete dictionary (more dictionary atoms than input dimensions) lets the SAE unpack the polysemantic neurons into many monosemantic features.- Vanilla L1 SAE (Bricken et al., 2023): the original. Suffers from "shrinkage" — L1 also penalizes correct nonzero values, biasing the reconstruction.
- TopK SAE (Gao et al., 2024, OpenAI): replace L1 with a hard top-
kmask — keep only theklargest activations, zero the rest. No shrinkage, cleaner sparsity-reconstruction tradeoff. - JumpReLU SAE (DeepMind, 2024): a learned activation threshold per feature, using a straight-through estimator. State of the art on the Gemma Scope suite.
- Gated SAE: separates the "should this feature fire?" decision from the "how strongly?" decision into two heads.
#Hands-on: A Toy Sparse Autoencoder
d_model = 4096 and d_dict = 16,777,216.#Hands-on: Build an Induction Head by Hand
Now let us build the canonical interpretability circuit — a two-layer attention pattern that performs A,B,...,A → predict B — without any training. Just hand-set weights.
b at the final position. Two simple operations — copy your left neighbor's identity, then match on identity and read out what came next — and you have the mechanism behind in-context learning. The real induction heads inside Claude or LLaMA do this with learned W_Q, W_K, W_V, W_O matrices that approximate the construction above. The algebra is the same.#Part 7: Activation Patching and Causal Intervention
The recipe (Vig et al. 2020 introduced it; Meng et al. 2022 popularized it as "causal tracing"):
- Pick a corrupted prompt that the model should get wrong (e.g., random tokens substituted for the key entity).
- Pick a clean prompt that the model gets right.
- Run both forward passes. Cache every hidden state from the clean run.
- For each (layer, position) pair, run the corrupted forward pass again, but at that one spot, replace the corrupted hidden state with the cached clean one.
- Measure how much the output recovers toward the clean answer.
What is the key difference between activation patching and path patching?
#Part 8: Logit Lens and Tuned Lens
nostalgebraist in 2020 and now a standard tool.The lens trick:
# At each layer L, compute as if we were finishing here:
logits_L = layer_norm_final(x_L) @ unembed.T
top_token_at_layer_L = logits_L.argmax(axis=-1)Plot the top predicted token as a function of layer depth. You will see things like:
- The model commits to its final answer many layers before the last layer (often around 60-80% of depth).
- Early layers' predictions are typically noisy or copy-the-input.
- For factual questions, the prediction tends to "snap" to the answer in a specific narrow range of layers — usually shortly after the MLP layer that ROME would identify as storing the fact.
#Part 9: Mech Interp as a Practical Debugger
This is not just academic. By 2026, interp tools are starting to be operationally useful:
- The chat model refuses harmless queries. Symptom: asks "How do I sharpen a kitchen knife?" → "I cannot help with that." Procedure: extract a refusal direction from refusal-vs-comply prompt pairs; check whether it fires on the harmless query; ablate it at the refusal-mediating layer; verify behavior changes. (Arditi et al.'s methodology, applied as debugging.)
- The model hallucinates dates. Symptom: confidently says JFK was assassinated in 1962. Procedure: causal-trace which MLP carries the date; either edit the MLP rank-one (ROME) or, for production, route around it with a retrieval system.
- The model confidently lies under pressure. This is still research, not solved. Burns et al.'s CCS, the geometry-of-truth direction, and SAE features for "deception" all suggest there is something there to find, but isolating a reliable, generalizable "I am about to deceive the user" feature is an open problem (Conmy et al., 2024; ongoing).
- Inserting watermarks or backdoors detection. SAEs have been shown to surface "training-data-poisoning style" features. Mech interp is the natural defense against worst-case model tampering: you cannot hide an algorithm from a sufficiently good circuit-finder.
Playground: Attention Heatmap — explore real attention patterns layer-by-layer and try to spot induction-head and name-mover-head signatures.
#Part 10: Open Problems and the Road Ahead
- Scaling SAE evaluation. We can train SAEs with millions of features. We cannot reliably evaluate which ones are real, which are training artifacts, and which are spuriously firing on tokenization patterns. The field needs better automated feature interpretability and reliability metrics.
- Feature composition into circuits. A 34M-feature SAE shows you the atoms. The compositional question — "which features combine, through which heads, to implement multi-step reasoning or planning?" — remains hard. The IOI-style circuit decomposition over SAE features (rather than over raw activations) is the active frontier.
- Causal sufficiency vs necessity. Most circuit analyses establish necessity (ablating breaks the behavior). Sufficiency (these components alone reproduce the behavior in isolation) is harder and often fails because of redundancy and backup components.
- Multimodal interp. Vision-language models, audio models, embodied agents. SAEs on Gemma image-text representations are new and the interpretive language is still being developed.
- Interpretability at inference, not just at training. If we want to use interp to monitor deployed models, the analysis has to run in real time, in-line with serving traffic. This is mostly an engineering problem (KV-cached SAE encoders, lightweight feature dashboards) but it is what turns interp from a research artifact into safety infrastructure.
- Adversarial robustness of mechanistic findings. Can a model be trained that looks interpretable but hides its real algorithm behind a Potemkin facade? Anthropic's "Sleeper Agents" work showed at least one form of hidden behavior survives safety training. Whether SAEs can be hidden from is open.
The bet, made loudly by Anthropic and quietly by every serious safety team: mechanistic interpretability scales. If it does, deployment of frontier models becomes a question of auditing internal features, not just behavioral red-teaming. If it does not — if the algorithms inside trillion-parameter models prove genuinely resistant to mechanical understanding — then we will be deploying systems that we operate but do not comprehend. The interpretability research of the next five years will determine which world we live in.
#Key Takeaways
- The residual stream is the substrate. Every feature in a transformer is, at some point, written into the residual stream by some head or MLP. Reverse-engineering a model is reverse-engineering reads and writes to that stream.
- Induction heads are the canonical circuit. A previous-token head plus a same-token-match head produces in-context copying, and the emergence of induction heads tracks the emergence of in-context learning.
- Polysemanticity is the rule, not the exception. Individual neurons are not features. Superposition forces meaning into linear combinations.
- Sparse autoencoders unpack superposition. L1 (or TopK / JumpReLU) regularized overcomplete dictionaries trained on frozen activations yield millions of monosemantic features — Anthropic Claude 3 Sonnet at 34M, DeepMind Gemma Scope publicly.
- Activation patching is the causal hammer. Replace clean hidden states into corrupted runs to localize where a behavior is computed. Path patching gives edge-level precision.
- Activation steering is mechanistic control. Once a feature direction is known, adding or subtracting it at inference time can amplify or suppress the corresponding behavior — including, alarmingly, RLHF-induced refusal.
- The logit lens is a transformer debugger. Project the residual stream through the unembedding at every layer to see when the model commits to its answer.
#Quick Check
The logit lens projects the residual stream at every intermediate layer through the unembedding matrix. What does it primarily tell you?