Circuits made of features rather than neurons

The circuit discovery work of the last two years has mostly drawn graphs whose nodes are attention heads, MLP layers, or individual neurons. Those units are what the architecture gives you, and they are also polysemantic, so the graphs explain a behaviour in terms of components that themselves need explaining. Samuel Marks, Can Rager, Eric Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller posted a paper on March 28 that swaps the nodes out. Their circuits are built from sparse autoencoder features, the dictionary directions that the monosemanticity line of work has been producing since last autumn.

For Pythia-70M they trained their own SAEs on every sublayer with a dictionary 64 times the model dimension. For Gemma-2-2B they used the public Gemma Scope dictionaries at 8 times the model dimension. Given a task, they estimate each feature's causal effect on the output with attribution methods, keep the features above a threshold, and connect them with edges weighted by their estimated influence on each other. The result is a graph where each node is something you can read, a feature that fires on plural nouns, or on female names, or on tokens inside a relative clause.

The agreement circuits

The first test is subject-verb agreement, the standard toy problem for this kind of work, in four syntactic frames: simple sentences, agreement inside a relative clause, agreement across a relative clause, and agreement across a prepositional phrase. The circuits are small. Roughly 100 feature nodes for Pythia-70M and roughly 500 for Gemma-2-2B recover half the model's performance on the task, against the thousands of neurons a neuron-level circuit would need for the same coverage.

The compression is the point. A 100 node graph is something a person can sit with and annotate. The authors do that, and the features they find in the circuits line up with the linguistic story you would expect, number information flowing from the subject through to the verb position. That was already believed. What is new is that the graph says it in vocabulary a linguist could check.

SHIFT: deleting the shortcut

The part of the paper we think will be cited most is the application. They call it Sparse Human-Interpretable Feature Trimming, SHIFT. The setup is the Bias in Bios dataset, where the task is to classify a biography as belonging to a professor or a nurse. In the training split the labels are correlated with gender, so a linear probe on the model's activations learns a mixture of profession and gender and does badly on a balanced test set where the correlation is broken.

SHIFT builds the feature circuit for the classifier, has a person look at each feature and mark whether it is relevant to the intended task, and ablates the ones that are not. The human sees features and their top activating examples, and judges that a feature which fires on she and her and female first names has nothing to do with whether someone is a professor. They marked about 55 such features in Pythia and 43 in Gemma.

The numbers on the balanced evaluation set are large. The Pythia-70M classifier starts at 61.9 percent profession accuracy with 87.4 percent gender leakage, meaning a probe can recover gender from what the classifier attends to almost nine times out of ten. After SHIFT it is 88.5 percent profession accuracy and 54.0 percent gender leakage, close to chance. Retraining the classifier after ablation reaches 93.1 percent and 52.0 percent. For Gemma-2-2B the sequence is 67.7 to 76.0 to 95.0 percent on profession, with leakage dropping from 81.9 to around 52 percent.

Why this is different from what came before

There have been methods for removing spurious correlations before, and some of them work well. What they have in common is that they need the spurious signal labelled in the data. You tell the method that gender is the nuisance variable and it projects it out or reweights against it. SHIFT does not use gender labels. The human is looking at the model's own features and deciding which ones belong to the task, using only the knowledge of what the task is supposed to be.

That is a meaningfully different workflow. It means the editor does not need to know in advance what shortcut the model took, only to recognise one when the circuit shows it. It also means the same procedure applies to a shortcut nobody anticipated, as long as the shortcut shows up as a readable feature. We think that is the first time dictionary features have been used to change what a model does for a downstream purpose rather than to describe it, and it is a much stronger argument for the SAE programme than another feature gallery.

What the small scale leaves open

The authors say much of their evaluation is qualitative, and they are right to say it. The debiasing demo is a linear probe on a 70M model and a 2B model, on a two-class task, with one well-understood nuisance variable. Nothing in the paper tells us whether SHIFT works when the shortcut is diffuse across hundreds of weak features, or when the human cannot tell from top activations whether a feature is task-relevant, or on a model where the relevant computation happens in a layer the dictionaries reconstruct badly.

There is also the cost. The method presupposes trained SAEs for every sublayer, which for Pythia they had to make themselves. That upfront compute is fine for a 70M model and becomes the dominant expense for anything larger. Gemma Scope removes it for one model family, and the fact that the Gemma numbers are as good as the Pythia numbers is the most encouraging detail in the paper, because it says the method survives someone else's dictionaries.

The paper also includes an unsupervised pipeline that clusters contexts from The Pile and builds circuits for thousands of automatically discovered behaviours, but that section is a demonstration of scale rather than a result anyone has checked. What we would like to see next is SHIFT applied to a classifier where the shortcut is not gender, chosen by someone who does not know what the shortcut is, and scored on whether they find it. That is the experiment that tests the claim rather than the demo.

Sources

  1. Marks, Rager, Michaud, Belinkov, Bau, Mueller, Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models (arXiv 2403.19647)