A team that built the tool stops building it

On 26 March the DeepMind mechanistic interpretability team, the same group that released Gemma Scope last August, published a progress update saying they are deprioritising fundamental sparse autoencoder research. The reason is a negative result. They picked a downstream task they thought sparse features should win, ran the comparison against the simplest possible baseline, and lost.

We want to be careful about tone here because the post is being read as a repudiation and it is not one. The authors say explicitly that they do not think SAEs are useless, and they list things the method remains good for. What they say is that after a year of heavy investment they could not find a compelling scenario where SAEs beat baselines, and that if the method were the step forward people hoped, that should not have been so hard.

That is the correct way to run a research programme and it is rarer than it should be. Most groups would have kept publishing incremental improvements on the sparsity-reconstruction curve. This group set a bar that other people cared about, measured against it, and reported the miss.

The experiment

The task was detecting harmful user intent in prompts, chosen because it is a safety problem with a clear label and because the interesting failure is generalisation. They trained on harmful prompts from HarmBench against harmless prompts from Alpaca, 4,046 training examples and 1,014 for validation. The out-of-distribution test set had 3,850 examples drawn from OverRefusalBench, internal chat data, and AdvBench, with jailbreak augmentations applied on top.

The distribution shifts were the kind you would meet in production. Adversarial ones like the AIM prompt, GCG suffixes, and roleplay framings. Semantic ones like word counting, instructions to ignore prior text, capitalisation changes, and synonym substitution. Training used the weaker transformations and testing used held-out variants, so a probe had to generalise across shift type, not just memorise the augmentations it saw.

The hypothesis going in was that a sparse, interpretable representation would be less sensitive to spurious correlations than a dense one, and so would hold up better under shift. The dense linear probe on raw activations reached an AUROC of 1.0 in training and validation and 0.999 out of distribution. A single SAE latent from the base dictionary managed roughly 0.60 out of distribution. Finetuning the latent brought it to about 0.85. A k-sparse probe over five to twenty finetuned SAE latents reached about 0.98 on validation but only 0.92 to 0.95 out of distribution.

Why a few points of AUROC matter

Ninety-five percent AUROC sounds fine until you remember that the baseline it lost to is a single matrix multiply that costs nothing to train and needs no dictionary. The whole argument for paying the SAE cost, in this setting, was that sparsity would buy generalisation. It bought the opposite. The sparse probes matched the dense probe in distribution and fell behind under shift, which is the reverse of the prediction.

The authors connect this to parallel negative results from other groups on probing, unlearning, and steering, and to a list of known problems with current dictionaries: incomplete coverage, noisy features, feature absorption, and poor sensitivity. Their reading is that current SAEs do not find the concepts required to be useful on an important task. Note the hedge. Current SAEs. The claim is about the dictionaries we can train today, not about the idea.

The line we keep coming back to is their speculation that the community is somewhat over-invested in SAEs. A method that became the default so quickly, across so many groups, on the strength of feature browsers and one very good bridge demo, was always going to need a correction. This is what the correction looks like from a group that has the compute to run the test properly.

What the post says SAEs are still for

The most concrete positive finding in the post is about dataset debugging. When the team looked at which SAE features fired on the harmful and harmless training sets, the features revealed spurious correlations, for example a feature that distinguished questions in the harmless set from those in the harmful set. That let them clean the data by hand before training the linear probe. So the interpretable representation was useful, but as a diagnostic step on the way to a baseline rather than as a replacement for it.

They also keep SAEs for exploratory analysis, meaning cases where you do not know in advance what feature you are looking for, and for debugging mysterious model behaviour. And they recommend that future SAE work focus on understanding fundamental limitations rather than hill-climbing on the sparsity-reconstruction trade-off. The team itself is moving toward model diffing, interpretability of deception in model organisms, and understanding reasoning models.

A position paper from Peng, Movva, Kleinberg, Pierson, and Garg, circulated since June, makes the same distinction as a thesis. SAEs are tools for discovering unknown concepts, not for acting on known ones. Probing, steering, and unlearning all start from a concept you already have, so a supervised method with that concept as its label will usually win. The discovery setting has no label to give a supervised method, and that is where an unsupervised dictionary has a job. We think this is right and we think it is what the DeepMind result shows rather than contradicts.

What we take from this

The practical rule we are adopting is simple. If you can write down the concept, train a probe. If you cannot, train a dictionary and use it to find out what the concept is, then train a probe. The dictionary is the discovery phase and the probe is the deployment phase, and confusing the two is how the field got a year of disappointing steering results.

What we would like someone to do next is build an evaluation for the discovery use case that is as clean as the OOD probing setup was for the acting use case. Plant a concept in a model that no one on the team knows about, hand the dictionary to a second team, and score whether they find it. Until we have that, claims that SAEs are good for discovery are as untested as the claims about generalisation were a year ago.

We are going to keep using the Gemma Scope dictionaries, but we are going to stop treating a feature as a finding. A feature is a lead. The finding is the probe you train once you know what to look for, and the ablation that shows the probe tracks something the model uses.

Sources

  1. Negative Results for SAEs On Downstream Tasks and Deprioritising SAE Research (Alignment Forum)
  2. Same post on the DeepMind Safety Research Medium
  3. Position: Use Sparse Autoencoders to Discover Unknowns