What was released

Two months after the attribution graph papers, the method is now something you can run. On May 29 Anthropic announced circuit-tracer, an open-source library written by Michael Hanna and Mateusz Piotrowski during their time as Anthropic Fellows, with Emmanuel Ameisen and Jack Lindsey as mentors. The library generates attribution graphs, the directed graphs of features that trace how a model gets from a prompt to a particular output token, on open-weight models. Gemma-2-2B and Llama-3.2-1B are the supported starting points.

The other half of the release is the front end. Johnny Lin and Curt Tigges at Decode Research, who run Neuronpedia, built an interactive graph explorer that hosts the graphs, lets you annotate nodes, and share a link to a specific graph. You can also intervene, setting feature values up or down and watching the output distribution change, which is the step that turns a graph from a picture into a hypothesis test. A tutorial notebook walks through multi-step reasoning and multilingual examples on both models.

What changes when a method leaves one lab

Until this week, attribution graphs were something only Anthropic could produce, on Claude models nobody else could inspect. The results in the March papers were interesting and there was no way to check them. Every claim about how the model computed a capital city or planned a rhyme had to be taken on the strength of the figures, and every question about whether the same structure exists in other models was unanswerable.

That is the situation the release addresses. The graphs on Gemma-2-2B are the same kind of object as the graphs on Claude, produced by the same procedure, on a model with public weights and public dictionaries. If the Dallas to Texas to Austin chain shows up in a 2B open model, that is evidence the finding was about language models rather than about one company's training recipe. If it does not, that is evidence too, and until now nobody could collect either kind.

There is a second effect that matters more for our kind of work. Attribution graphs require transcoders or SAEs for every layer, plus the machinery to compute attributions through them. Building that from the papers alone would have taken any outside group months. Having it as a pip install moves the bottleneck from engineering to the question of what to look at, which is where a research field wants its bottleneck to be.

What the examples show

The worked examples in the release are the ones from the original papers, reproduced on the small models. The multi-step reasoning case is the capital of the state containing Dallas, where the graph shows a Texas feature active between the prompt and the Austin output. The multilingual case shows features that fire regardless of the language of the prompt sitting alongside features specific to one language, which is the structure the March paper reported on Claude.

What the release does not include is a systematic comparison between the graphs on the small open models and the published graphs on Claude. The examples are chosen to match, and they do match, which is the point of a demonstration. Whether the match holds across behaviours nobody has selected for is the question the tool now makes answerable, and it is the one we would want answered before treating the Claude results as general facts about language models.

What the tool does not do

The library paper is direct about the limits and so was the coverage. Attribution graphs are expensive in memory, since you hold the model, the dictionaries, and the attribution computation at once, and the two supported models are small partly for that reason. The graphs are also hard to read, and the front end helps with navigation without helping with interpretation. Whether a given feature means what its top activations suggest is still a judgment the user makes.

There is a deeper caveat that the release does not resolve. Attribution graphs are built through replacement models, dictionaries that approximate the real MLP computation, and the graph describes the replacement rather than the original. The replacement is an approximation with error that the papers report rather than eliminate. A circuit found in the replacement is a candidate that has to be confirmed by intervening on the real model, which the tool supports, and which most people using it for the first time will skip.

What we would want outside users to try

The obvious first project is replication. Take the behaviours reported on Claude, run them on Gemma and Llama, and publish a table of which ones appear, which ones appear differently, and which ones do not appear at all. That is unglamorous and it is what the field needs, because the March papers are currently a single-lab result on a closed model.

The second project is to pick a behaviour nobody has traced, something small and well-defined, and find out whether the graph tells you anything you could not have got from a probe. If the answer is usually yes, the method earns its cost. If the answer is usually that the graph confirms what a simpler tool would have shown, that is worth knowing before more labs invest in building dictionaries for every layer of every model they own.

Sources

  1. Anthropic, Open-sourcing circuit tracing tools (May 29, 2025)
  2. Hanna, Piotrowski, Lindsey, Ameisen, Circuit-Tracer: A New Library for Finding Feature Circuits (BlackboxNLP 2025)
  3. VentureBeat, coverage of the circuit tracing release