Why test a method on a solved problem

Anthropic's circuit tracing work, published on March 27, built attribution graphs for prompts on a production model and reported mechanisms for poetry planning, multi-step recall and refusals. Those findings were new, which is what made them exciting and what made them hard to check. Nobody knew the ground truth for how Claude 3.5 Haiku plans a rhyme, so nobody could say whether the graph was right.

On June 11 Goodfire published a replication with a different goal. Jack Merullo, Connor Watts, Max Loeffler, Liv Gorton, Elana Simon, Tom McGrath and Owen Lewis applied the same pipeline to a mechanism that has been mapped by hand. If the graphs recover what we already know, that is evidence for trusting them where we do not. If they miss it, that is a limit worth knowing before anyone builds a safety argument on a graph.

The mechanism

The target is the greater-than task from Hanna, Liu and Variengien's 2023 paper. Give GPT-2 Small a prompt like The war lasted from the year 1732 to the year 17 and it puts most of its probability on two-digit completions above 32. The original work traced this to a circuit ending in the later MLP layers, which boost end years greater than the start year, and showed the behaviour holds across many surface contexts. It is a small, clean, verified piece of model internals, which makes it close to ideal as a calibration target.

The pipeline

Attribution graphs depend on a replacement model. The MLP layers are swapped for cross-layer transcoders, sparse feature dictionaries where a feature reads from the residual stream at one layer and writes to every later layer. On a given prompt, attention patterns are frozen, error terms are added for whatever the transcoders fail to reconstruct, and the interactions between features become linear, so you can draw a graph with features as nodes and direct effects as edges. Anthropic's own methods paper lists the limitations: attention pattern formation is not explained, reconstruction is partial, and inhibitory interactions are poorly captured.

Goodfire trained their transcoders on 100 million tokens of FineWeb, with 49,152 features per layer across GPT-2 Small's 12 layers, 589,824 features in total. The transcoder reached an L0 of 79 and 59 percent token accuracy, meaning the replacement model reproduces the original's top prediction on 59 percent of tokens. That number is worth pausing on. Four in ten tokens, the replacement model and the real model disagree about the next word, and the graph is a picture of the replacement.

What the graphs recovered

They found the mechanism. Among 160 features of interest before deduplication, there are greater-than features that push the output toward larger years, matching the role Hanna and colleagues assigned to the late MLPs. There is a parity feature that promotes even-numbered years and suppresses odd ones, which is the kind of structure a hand analysis is unlikely to notice. There are structured-number features responding to multiples of five and to sequential patterns. The graph produces a story that agrees with the 2023 paper at the level of what the circuit does and where.

The disagreements are where the paper earns its place. The authors describe the representation space as less clean than Hanna's account. Some features fire on extremely narrow input ranges, a handful of specific years rather than a smooth greater-than direction. The structured-number features have no counterpart in the original circuit, which either means the original missed them or means the transcoder invented a decomposition that the model does not use. And there are discrepancies between the logit attribution heatmaps, which say what a feature does to the output, and the max-activating examples, which say when it fires. A feature that looks like a greater-than detector by its output effect can have activating examples that suggest something else, which the authors read as features interacting contextually rather than acting alone.

What this buys the rest of us

The honest summary is that the pipeline passed, with caveats that would be invisible on an unsolved problem. On a mechanism where we know the answer, the graph found it and also found extra structure of uncertain status. On Haiku-sized mysteries the extra structure would be reported as discovery, because there is nothing to check it against. That is the value of the exercise. It puts a rough error bar on the method that you cannot get from the exciting results.

Interpretability has run for years on findings that were interesting and unverified, and a replication on GPT-2 Small will never trend. It should still be the default first step for any new method. We would like to see the same calibration done on the other hand-mapped circuits, indirect object identification and docstring completion at least, with the transcoder's reconstruction accuracy reported next to every graph. When a method recovers all three with the same caveats, the caveats become a known property of the tool. Until then each new attribution graph on a frontier model should be read with the 59 percent in mind.

Sources

  1. Goodfire: Replicating circuit tracing for a simple, known mechanism
  2. Anthropic: Circuit Tracing, revealing computational graphs in language models
  3. Hanna, Liu, Variengien: How does GPT-2 compute greater-than? (arXiv 2305.00586)