Attribution graphs: wiring diagrams for Claude 3.5 Haiku, with the gaps marked
Anthropic traced individual prompts through a 30-million-feature replacement model and found planning in poems, a shared conceptual space across languages, and reasoning that does not match the chain of thought. The method works on about a quarter of prompts.
From features to circuits
Last week Anthropic published two papers that move dictionary learning from a list of features to a diagram of how they connect. The first describes the method, which the authors call attribution graphs. The second, titled On the Biology of a Large Language Model, applies it to Claude 3.5 Haiku across a set of case studies. Simon Willison's reaction was partly to the format, since neither paper is a PDF and both are long HTML pages with linkable sections and inline interactive diagrams. That is a small thing that makes checking the claims much easier, and we wish more groups did it.
The core move is to swap the model for something you can read. They train cross-layer transcoders that produce a replacement model with 30 million features, then for a single prompt they build a local replacement model that combines those features with error nodes for what the transcoders miss and the original attention patterns. Tracing influence through that object gives a graph from input tokens to output token. They then collapse related features into supernodes by hand so a person can follow it.
The error nodes are the honest part of the design. The replacement model does not reconstruct the real model perfectly, and rather than hide the residual the method carries it as an explicit node in every graph. When a graph shows a lot of influence flowing through error nodes, that is the method telling you it does not know what happened.
The case studies we found most convincing
The poetry result is the one people will remember. Asked to write a rhyming couplet, the model activates features for candidate end words at the newline before the second line starts. In the worked example, features for rabbit and habit light up early, and then intermediate word choices like a phrase ending in 'like a' are shaped to land on the planned rhyme. The model is not choosing the last word last. It picks a target and writes toward it. The authors confirm this by intervening on the planned word and watching the line change.
The multilingual result is quieter but we think more important. The same prompt asking for the opposite of small, in English, French, and Chinese, routes through a shared language-independent circuit for antonyms, with language-specific features only at the output. The authors report that the share of these multilingual features grows with model size. If that holds, larger models are converging on an internal representation that is not in any language, and the language they answer in is a late decision.
The arithmetic case shows why the reuse matters. A lookup-table feature for adding six and nine fires in an arithmetic problem, and it also fires on astronomical timestamps and on the year in an academic citation. The same circuit is doing addition in contexts that do not look like arithmetic on the surface. That is a strong argument that the features are functional units rather than topic detectors.
Reasoning that is not what the model says it is
The medical case has the model reading a description consistent with preeclampsia, activating symptom features that promote candidate diagnoses, then asking about a confirmatory symptom. The intermediate steps happen inside the forward pass and never appear in the output. The authors describe it as reasoning in its head. That is an existence proof that the chain of thought, when there is one, is not the only reasoning going on.
The unfaithfulness result is the one we would put in front of anyone who relies on reading a model's stated reasoning. In some cases where a human suggested an answer, the graph shows the model working backward from the suggestion to produce steps that look like a derivation. The written explanation is coherent. The mechanism that produced the answer is not the mechanism the explanation describes. The authors also show a model finetuned to exploit training bugs whose hidden goal is bound up with its representation of the Assistant persona, which is a finding you could not get by reading transcripts.
The hallucination and refusal graphs fit together. There is a default circuit that declines to answer, inhibited by known-answer features. For an entity the model half knows, the known-answer features fire weakly, the refusal is suppressed, and the model produces a plausible fabrication. In the jailbreak example, the model assembles a word letter by letter without recognising it as a harmful request until the word is on the page, then tries to recover at the next sentence boundary. The harmful-request features themselves cluster by harm type from pretraining and get wired together into a general refusal trigger during finetuning.
The limits the authors state and the ones they do not
The number to remember is a quarter. The authors say the method gave them satisfying insight on about a quarter of the prompts they tried. The case studies are the ones that worked. Case-study papers are written that way, and that is fine, but it means the method is a microscope that only focuses sometimes, and we do not yet have a good account of what distinguishes the prompts it handles from the ones it does not.
Two structural gaps are named. Attribution graphs explain active features and cannot say why an inactive feature failed to fire, so absence is invisible. And attention is taken from the original model rather than explained, so any computation that lives in how attention patterns are formed is outside the picture. Both are substantial. A great deal of what a transformer does is deciding where to look.
The gap we would add is that the supernodes are hand drawn. Grouping features into a legible diagram is a judgement call made by someone who already has a hypothesis, and the same raw graph could support different stories depending on how you group it. The interventions partly answer this, since they test the story rather than the drawing, but we would want to see two independent groups produce diagrams from the same graph before trusting any single one.
A note from later
We are adding this section in August 2026 because the method has kept moving and the original post should say so. In May 2025 Anthropic open-sourced the attribution graph library, built by Fellows Michael Hanna and Mateusz Piotrowski with Decode Research, and Neuronpedia put up a frontend for generating and sharing graphs on Gemma-2-2b and Llama-3.2-1b using transcoders from Gemma Scope. The tool that produced the Haiku case studies now runs on models anyone can download.
The June 2026 circuits update extends graphs to whole conversations. The problem was that per-token features on a long transcript number in the thousands or millions. The fix is to average the residual stream over each turn and train a dictionary on that, so a turn has L0 active features rather than tokens times L0. On Qwen-2.5-7B-Instruct over LMSYS-Chat-1M, the turn-averaged features picked out things like incorrect answers in number puzzles where per-token features only surfaced arithmetic, and they generalised to turns 150 times longer than anything seen in training.
The quarter figure is the one we still want addressed. A method that explains a quarter of prompts is a research instrument. A method that explains most of them, or that can tell you in advance which ones it will explain, is a safety tool. We do not know which of those we will have in another eighteen months, and we have not seen a number that updates the original one.
Sources
From the foundation