What was published

Last week Anthropic's interpretability team released a paper showing that a single MLP layer with 512 neurons in a small transformer can be decomposed into more than 4,000 features, most of which a human can label. The method is a sparse autoencoder, which the authors frame as dictionary learning. You train a second, wider network to reconstruct the layer's activations through a sparse bottleneck, and the directions the bottleneck learns are the features.

The examples in the paper are the kind you can check yourself. There are features for DNA sequences, for legal language, for HTTP requests, for Hebrew text, and for nutrition statements on food packaging. The authors are blunt that most of these properties are invisible if you look at individual neurons. A single neuron in the layer fires on a mixture of unrelated things. The features do not.

We have read the paper twice now and spent an evening with the interactive feature browser. This post is our attempt to say what we think the result establishes, where the open questions sit, and why we expect people to be building on it for a while.

Why the neuron was never the right unit

The background is the 2022 toy-models-of-superposition work, which argued that a network with more useful concepts than dimensions will pack several concepts into overlapping directions. If that is true, then asking what a neuron means is asking the wrong question, because the neuron is a basis vector the optimizer never cared about. The concepts live in directions that are not axis aligned, and a neuron will respond to whichever of them happen to project onto it.

That was a compelling story, but until now it was mostly a story about toy models. What this paper adds is a procedure that recovers those directions from a real (if small) language model and a demonstration that the recovered directions look like concepts. The word the authors use is monosemantic. One feature, one meaning. The neuron by contrast is polysemantic, and the paper shows this by direct comparison.

The comparison is worth dwelling on because it is the closest thing in the paper to a measurement. Human raters scored how interpretable a unit was. On the scale they used, the median feature scored 12. The median neuron scored 0. We would like to see that experiment replicated by a group with different raters and a different rubric, but a gap that size does not go away with a rubric change.

What the features can do beyond being labelled

A label on a feature is a hypothesis. The paper does two things to test the hypotheses. First, it checks that features are predictive, meaning that the activation of the feature tracks the presence of the concept in the input. Second, it intervenes. Artificially activating a feature causes the model's output to change in the direction the label predicts. If you turn up a feature that responds to a particular kind of text, the model starts producing that kind of text.

This second test matters more than the first. Correlational interpretability has a long history of finding patterns that turn out to be epiphenomenal. Causal interventions are harder to fool. They are still not proof that the feature is what the model uses internally, because a direction can be causally effective without being the natural unit of computation, but they rule out the most embarrassing failure mode.

The paper also reports that the number of features you ask the autoencoder to learn acts as a resolution control. Ask for fewer and you get coarser concepts. Ask for more and a coarse concept splits into finer ones. The authors describe this as a knob for varying the resolution at which you view the model. That framing is useful and it also hints at the main unresolved problem, which is that there is no principled way to pick the dictionary size.

The criticisms we find most serious

The discussion thread on the Alignment Forum has two objections we think the authors will need to answer. The first, raised by nostalgebraist, asks why the autoencoder was trained on MLP activations rather than on the residual stream update. If features are best understood as directions in the space of logit updates, then decomposing the input side of the layer might be capturing detectors rather than the things the model actually does with them. The paper does not settle this.

The second, from Tom Angsten, points out that superposition should not in theory help the final MLP layer of a one-layer model, because there are no downstream nonlinearities to exploit the packed representation. And yet the dictionary keeps improving past 512 features. Either the theory is incomplete or the autoencoder is finding structure that is not superposition in the original sense. Both would be interesting. Neither is resolved.

Our own concern is more mundane. The model is tiny. The paper is explicit that it applies to small transformer models, and there is a real chance that what looks clean at 512 neurons becomes a mess at frontier scale, where a layer might have tens of thousands of neurons and the number of concepts is unknown. The authors seem aware of this. Their closing line is that the next primary obstacle to interpreting large language models is engineering rather than science. That is a strong claim and we hope they are right, but it is a claim about work not yet done.

What we would try next

The obvious experiment is to run the same method on a model people actually use, with a dictionary large enough to matter, and report how many features survive human inspection. We would also want a test that does not depend on human raters, since the raters know which condition they are in. Automated interpretability scoring, where one model writes an explanation and another predicts activations from it, is the natural candidate and the paper already uses it in places.

The second thing we would try is the residual stream question directly. Train the autoencoder on the residual stream at the same depth, compare the dictionaries, and see whether the features that look like detectors on the MLP side turn into features that look like actions on the stream side. If the two dictionaries are largely the same, the objection dissolves. If they differ, we learn something about what the layer is for.

For a non-profit like ours the appeal of this method is that it is cheap. A sparse autoencoder on a small model trains on a single GPU, and the code path is short. We are going to reproduce the headline numbers on an open model over the next month and we will write up whatever we find, including the parts that do not match.

Sources

  1. Towards Monosemanticity: Decomposing Language Models With Dictionary Learning
  2. Anthropic research summary
  3. Alignment Forum discussion thread