The question nobody had checked

Almost every circuit paper studies a single model at a single moment, the end of pre-training. Deployed models are not like that. They get trained further, fine-tuned, and replaced by larger siblings, and it has never been clear whether a circuit found in GPT-2 Small says anything about what happens after another hundred billion tokens or in a model forty times bigger. Curt Tigges, Michael Hanna, Qinan Yu and Stella Biderman put out a paper on July 15 that asks this directly, and the answer is more reassuring than we expected.

The setup uses the Pythia suite, which is the only public family with checkpoints at many scales. Each model has 154 checkpoints spanning 300 billion tokens, and the paper finds circuits at every one of them for models from 70 million to 2.8 billion parameters, with behavioural data up to 12 billion. The tasks are four staples of the circuits literature: indirect object identification, greater-than, subject-verb agreement and gendered pronoun completion.

How they found circuits at 154 checkpoints

Patching-based circuit discovery needs a number of forward passes that grows with model size, which rules it out for this many checkpoints. The authors use edge attribution patching with integrated gradients, which scores every edge in the model's computational graph in a fixed number of passes, then binary search for the smallest set of edges that recovers at least 80 percent of the full model's task performance, starting from a search space of one edge to 5 percent of all edges. Faithfulness is checked by corrupting everything outside the circuit and seeing whether behaviour holds.

The 80 percent threshold is doing a lot of work here and the paper is candid about it. It buys confidence that the circuit carries most of the mechanism and gives up any claim to completeness, since no accepted method exists for showing a circuit contains every relevant edge. Everything that follows should be read as being about the main body of each mechanism rather than the whole of it.

Abilities and their components arrive on the same schedule

The first result is behavioural. Across all four tasks, models of different sizes learn the task at about the same number of tokens seen, and for each task there is a size beyond which more scale does not speed up learning and sometimes slows it. On IOI the 410 million to 2.8 billion models learn fastest, while the 6.9 and 12 billion models look more like the 160 million one. The authors say they found this surprising given other work showing larger models learning faster.

The second result explains the first. The attention heads known to carry these tasks emerge on the same schedule. Induction heads appear in most models soon after 2 billion tokens, replicating the earlier finding, and greater-than behaviour appears immediately after, along with successor heads. Name-mover heads for IOI appear between 2 and 8 billion tokens across scales, during or just before IOI performance rises, and copy suppression heads on the same timescale at varying strength. The ceiling on how fast these heads form looks like the ceiling on how fast the task is learned.

Same algorithm, different heads

The finding we would put on a slide is the one about what happens after formation. The strength of a given functional behaviour in a given head can fall later in training even while task performance holds steady. Zooming in on Pythia-160m, the paper tracks the top five heads for each role over time and shows the roster changing while the algorithm they implement stays put. The heads doing the name moving midway through training are not always the heads doing it at 300 billion tokens, and the circuit still moves the name.

There is also a Pythia-specific wrinkle worth knowing. In the original GPT-2 IOI circuit, copy suppression heads hurt the task by pushing down the correct name. In Pythia they help, pushing down the incorrect one, because both names are already strongly predicted by the time the signal reaches those heads and the mechanism suppresses whichever is more repeated. Same component type, opposite sign, same overall algorithm. Circuit size, meanwhile, grows with model scale and can swing considerably over training even when the algorithm does not change.

Why this is a quiet but essential result

The mechanistic interpretability programme rests on an assumption that almost never gets tested, which is that what you learn on a small open model tells you something about a large one. This paper is the first evidence at 300 billion tokens and across a forty-fold range of scale that the assumption holds for simple circuits. It holds at the level of algorithms and component types, and it fails at the level of individual heads, which is the right place for it to fail if you want to reuse findings rather than head indices.

The authors are careful to say this is for four simple tasks and for pre-training only, with no fine-tuning and nothing based on sparse autoencoder latents. Those are the two places we would push next. If a circuit for a harder task, or one expressed in SAE features rather than heads, turns out to be equally stable, then interpretability results on small open models start to look like genuine science about large closed ones. If it does not, we will at least know where the transfer stops.

Sources

  1. Tigges, Hanna, Yu and Biderman, LLM Circuit Analyses Are Consistent Across Training and Scale (arXiv 2407.10827)