The problem being solved

Neurons in language models are polysemantic. The same unit fires on academic citations, Korean text and HTTP requests, and no single description covers it. The explanation that has gained ground over the past year is superposition. The model represents more features than it has dimensions, so features share directions, and the neuron basis is the wrong place to look. Anthropic's toy models paper made that case in a synthetic setting. The open question was how to recover the features from a real model.

A paper posted this month by Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben and Lee Sharkey answers it with a method that is almost embarrassingly simple. Train a sparse autoencoder on activations. The encoder maps a residual stream vector to a dictionary that is larger than the model dimension, with an L1 penalty pushing most entries to zero. The decoder reconstructs the activation from the few entries that remain. The dictionary directions are the candidate features.

What they did

The models are Pythia-70M and Pythia-410M, small enough to train dictionaries on quickly. The dictionary size is set as a ratio R to the model dimension, and the sparsity penalty alpha was tuned per site, with values like 0.00086 and 3.2 times 10 to the minus 4 quoted for residual stream and MLP runs. The autoencoders are trained on activations gathered from a text corpus with no supervision about what a feature should be.

Interpretability is scored automatically, using the method OpenAI published in May for explaining neurons. A language model reads examples where a dictionary feature is active, writes a description, and then predicts the feature's activation on new text from that description alone. The correlation between predicted and real activations is the score. The same pipeline is run on neurons, on PCA components, on ICA components and on random directions, which gives a like-for-like comparison.

What they found

Dictionary features score far higher than any of the baselines, and the gap is largest in early layers. That is the headline result and it lines up with the theory. If features are packed into a residual stream in superposition, an unsupervised sparse decomposition should pull them apart, and PCA, which maximises variance rather than sparsity, should not. Neurons sit between, because an MLP's nonlinearity imposes some sparsity, and random directions come last.

The second result is causal rather than descriptive. On the indirect object identification task, the one where a model has to complete a sentence with the right name, patching a small number of dictionary features achieves the same edit as patching a larger number of PCA components. Features found by sparsity are a finer instrument for intervention as well as a cleaner one for description. That matters because an interpretability method that only produces nice labels is a curiosity. One that localises a behaviour to a handful of directions is a tool.

Why the convergence matters

The interesting fact about this paper is who wrote it. None of the five authors is at the lab that proposed superposition, and the work came out of the open interpretability community that has grown up around small models and shared code. They started from the same hypothesis, that the features live in a sparse overcomplete basis, and arrived at the same fix, dictionary learning, without a shared codebase or a shared employer.

That is what a hypothesis producing a method looks like. When one group proposes a mechanism and a different group, working from the mechanism alone, builds a technique that does what the mechanism predicts, the mechanism has earned some trust. We expect the large labs to publish their own dictionary learning results on their own models, and when they do, this paper is the reason the result will already be believable.

The caveats are the ones you would expect. Pythia-70M is tiny and nobody knows yet whether the feature count or the feature quality survives scaling. The autointerpretability score measures whether a language model can describe a feature, which is a proxy for whether a human can, and the OpenAI paper it borrows from was frank that individual scores are noisy. The dictionary size and the sparsity penalty are hyperparameters that trade reconstruction against interpretability, and the right operating point is a judgement call.

What to try

The obvious next step is scale, and the next obvious step after that is completeness. How much of a residual stream's variance is explained by interpretable features, and what is in the remainder. If the unexplained part shrinks as the dictionary grows, we are decomposing the model. If it plateaus, there is structure the method cannot see.

The experiment we would run first is cheaper. Take the features that score highest on Pythia-410M, find the nearest equivalents on Pythia-70M by activation pattern, and check whether the same features exist at both sizes. If the feature set is stable across models trained on the same data, the dictionaries are finding something about the data. If it is not, they are finding something about the model, and that is a different kind of object to study.

Sources

  1. Sparse Autoencoders Find Highly Interpretable Features in Language Models (Cunningham, Ewart, Riggs, Huben, Sharkey, 2023)