Does circuit analysis scale? DeepMind tries it on Chinchilla
Lieberum and colleagues found the multiple-choice circuit in a 70B model with the same logit attribution and activation patching that worked on GPT-2 small, but the semantics of the heads resisted a clean story. This is the first honest data point on how far toy-model methods carry.
The question in the title
Until this month, every circuit anyone had found in a language model lived in a small one. The indirect object identification circuit was in the 117M-parameter GPT-2 small, and the modular addition and docstring circuits were in models smaller than that. So the obvious worry was that the methods only work where the model is small enough to stare at. Tom Lieberum and six colleagues at Google DeepMind posted a paper on July 18 that takes the worry literally. They ran the standard toolkit on Chinchilla 70B, which has 80 layers with 64 attention heads each, and asked whether it still finds anything.
The task is multiple-choice question answering on MMLU, narrowed to the algorithmic part. Given that the model knows the correct answer text, how does it produce the letter A, B, C or D that labels it? The authors chose this because it is algorithmic, like the tasks where circuit analysis had worked before, and because it only shows up at scale. On five-shot MMLU, Chinchilla 1B and 7B score 25 and 26 percent when scored by label, while 70B scores 68. The 7B model scores 32 percent when scored by answer text, so it knows some answers and cannot do the symbol manipulation.
The part that scaled
The methods themselves carried over without modification. Logit attribution, meaning the direct effect of each head and MLP on the final logits, is computable for every component in parallel because the logits are a sum of component outputs given a fixed final norm. Averaged over 128 prompts, the direct effects had a few large contributors and a long tail. Forty-five nodes, 32 attention heads and 13 MLPs, explained 80 percent of the summed positive direct effect. Patching all 45 together from a prompt with a different correct answer recovered most of the loss and accuracy on the MMLU subset, which is the standard validation and it passed.
Attention pattern visualisation then sorted the heads into four groups. Correct letter heads attend from the final position to the label of the correct answer. Uniform heads attend evenly across all the labels. Single letter heads mostly attend to one fixed letter. Amplification heads, late in the network, appear to boost information already in the residual stream. Recursing one step further back, the authors found what they call content gatherer heads, with L24 H18 as the clearest example, which move information from the correct answer's content tokens to the final position so that the correct letter heads can use it in their queries.
Two things in this part deserve more attention than the summary gives them. The two nodes with the highest direct effect had noticeably lower total effect, and the authors have no satisfying explanation for the gap. And the per-letter breakdown of total effects was, in their words, somewhat confusing and at times contradictory, which they suspect is backup behaviour of the kind seen in the GPT-2 work, where the model compensates when a single node is patched.
The part that did not
Having found the correct letter heads, the authors tried to say what they compute. Using SVD on queries and keys collected from 1,024 prompts, they compressed each head's QK circuit to a three-dimensional subspace, which captured 65 to 80 percent of the variance for keys and queries, with a visible knee at three components, and 80 to 90 percent for values. Patching in the low-rank attention had the same effect as full-rank attention on the MMLU distribution. In that subspace the keys for the four labels form a tetrahedron and the queries cluster near the key for the correct answer.
Then they mutated the prompts. Changing the separators or removing the prelude made no significant difference, so the feature is not about formatting. Replacing A, B, C, D with random capital letters hurt, but the low-rank subspace still recovered a third to a half of the loss, so part of the feature generalises across letters and part is tied to those four specific tokens. Replacing the letters with 1, 2, 3, 4 broke the task entirely, even in the base model, and the correct letter heads appeared not to contribute at all in that setting. The authors' first hypothesis, that the subspace encodes Nth item in an enumeration, ended up as only a partial explanation.
Their pseudocode for the heads makes the problem visible. It names the keys item_nums and the query correct_item_num, and the text underneath says the names are only a first approximation, because the embedding for the second item points in roughly the same direction whether the label is B or a random letter, but with smaller magnitude, and smaller still for numbers. None of that fits in a line of code.
Why this is the honest data point
It would have been easy to publish the first half and stop. Existing techniques scale to 70B, here is the circuit, here is the diagram with content gatherers feeding correct letter heads feeding amplification heads. Instead the paper's abstract says the results on semantics are mixed, the discussion lists the caveats of resampling ablations at length, and the conclusion describes the results as relatively noisy and at times contradictory.
What we take from it is a split verdict that we think will hold. The mechanical part of circuit analysis, finding which nodes matter and validating them by patching, scales fine, and the cost is mostly engineering and labour. The semantic part, saying what a node computes in words that survive a change of distribution, was hard at 117M parameters and is not easier at 70B. The authors say progress on information flow has been fast and progress on what information is being processed has been comparatively slow, and this paper is the clearest demonstration of that gap so far.
What we would do next
The authors call the manual work very labour intensive and ask for automation, and we agree, but we would automate the boring half first. Finding the 45 nodes and validating them by patching is a procedure, and a procedure can be run on every algorithmic task in a benchmark rather than one. The interesting follow-up is to do that for a dozen tasks on the same model and see whether the same heads keep appearing. If the correct letter heads are reused for other enumeration tasks, the Nth item story gets stronger.
The other test is whether different models implement the same algorithm, which the paper lists as open. Running the same pipeline on an open 65B or 70B model would tell us whether the content gatherer and correct letter structure is a fact about multiple-choice answering or a fact about this model, and that is the difference between an interpretability result and a case study.
Sources
From the foundation