What launched

On December 11 a group spanning several interpretability teams, including Adam Karvonen, Can Rager, Johnny Lin, Curt Tigges, Joseph Bloom, David Chanin, Arthur Conmy, Callum McDougall, Samuel Marks and Neel Nanda, published SAEBench. It is a suite of eight evaluations for sparse autoencoders, plus results for more than 200 trained SAEs varying in sparsity, dictionary size, architecture and training duration, on Pythia-70M and Gemma-2-2B. The numbers are browsable on Neuronpedia.

For the past year the way you showed an SAE was good was a single curve: reconstruction loss against L0, the average number of active latents per token. A new architecture that pushed that frontier down and to the left got a paper. SAEBench's authors state the problem with this directly. The best SAE varies depending on the specifics of the downstream task, and the sparsity proxy has known failure modes, feature absorption and feature composition among them, that the curve cannot see.

The eight evaluations

Two of the eight are the unsupervised metrics people already used. Loss recovered measures how much of the model's original next-token loss survives when you swap the residual stream for the SAE's reconstruction. Automated interpretability has a language model judge whether each latent's top activations describe a coherent concept.

The remaining six are downstream tasks. Sparse probing asks whether a small number of latents can recover a labelled concept such as sentiment or language. Feature absorption checks whether a general concept, mammal say, has quietly stopped firing on tokens that a more specific latent, pig say, has absorbed. Spurious correlation removal takes a probe trained on a biased dataset and tests whether ablating a handful of latents removes the unwanted correlation. Targeted probe perturbation asks whether ablating latents for one class hurts only that class. Unlearning tests whether you can suppress a body of knowledge, biology facts in the paper's case, without damaging unrelated performance. RAVEL tests whether intervening on an entity's attribute changes that attribute and nothing else.

Each of those is a proxy for a thing someone might use an SAE for. That is the point of the design. The suite tells you whether the features will do a job, which is a narrower and more useful question than whether an SAE is interpretable in the abstract.

What the results say

The headline finding is that the sparsity-fidelity frontier does not reliably indicate performance on downstream tasks. Across the architectures tested, standard ReLU, Gated, TopK, BatchTopK, JumpReLU, P-Annealing and Matryoshka BatchTopK, an SAE that reconstructs better is not consistently one that probes better, unlearns better, or disentangles better.

The second finding concerns scale. As dictionary size grows from 4k to 16k to 65k latents, absorption gets worse for every architecture except Matryoshka. Bigger dictionaries split concepts more finely, and a finely split dictionary is one where a specific latent can steal activations from a general one. Matryoshka SAEs, which are trained so that nested prefixes of the dictionary must each reconstruct on their own, hold absorption steady as they scale.

Matryoshka is also the clearest example of the proxy metric misleading. It is slightly worse on reconstruction than the alternatives and substantially better on absorption, RAVEL, spurious correlation removal and targeted probe perturbation, with the gap widening at larger sizes. Under the old scoreboard it would have looked like a step backwards.

On sparsity the picture is a trade. Low L0 helps interpretability and hurts reconstruction, high L0 does the reverse, and the authors find a moderate range, roughly 50 to 150 active latents, that balances the metrics. They decline to collapse the eight numbers into one score, and we think they are right to. A single aggregate would just become the next proxy.

Why this changes what a paper has to show

Until this week, SAE research was largely an architecture race scored on a curve that nobody outside the field cared about. A new activation function or a new sparsity penalty could be evaluated in an afternoon and published on a plot. SAEBench makes that plot insufficient. If your architecture wins on loss recovered and loses on absorption and unlearning, reviewers can now ask which of those matters for the use you have in mind.

The practical effect is to move the question from 'is this a better autoencoder' to 'what are the features for'. A team building a probe for deceptive behaviour cares about sparse probing and spurious correlation removal. A team trying to remove hazardous knowledge cares about unlearning. Those teams can now pick an SAE on the metric that matches their job rather than on a general claim of quality.

It also gives a concrete home to negative results. Absorption was already a known failure mode. Now it is a column in a table for every SAE the suite has run on, and an architecture that ignores it will show up as ignoring it.

What we would want next

The obvious gap is model size. Pythia-70M and Gemma-2-2B are small, and the interesting claims about SAEs concern models people actually deploy. The suite is open, so the first thing we would try is running the six downstream evals on the largest SAEs anyone has released and checking whether the Matryoshka advantage survives.

The second gap is that six of the eight evaluations are still proxies, just better ones. Unlearning biology facts is a stand-in for unlearning something dangerous. Removing a spurious correlation from a profession classifier is a stand-in for debiasing a real system. The suite is a large improvement on one curve, and it would be a mistake to treat it as the final word on what features are for. We would like to see at least one evaluation where the SAE has to earn its place inside a real audit rather than a benchmark of one.

Sources

  1. SAEBench: A Comprehensive Benchmark for Sparse Autoencoders (LessWrong, December 11, 2024)
  2. SAE Bench on Neuronpedia
  3. Karvonen et al., SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability (arXiv 2503.09532, later paper version)