Automated interpretability for millions of features, on open models
EleutherAI's new pipeline explains sparse autoencoder latents with open models and scores the explanations for a few hundred dollars per million features instead of tens of thousands. Notes on what changed for labs without a frontier API budget, and on how the choice of scorer decides which features count as interpretable.
The cost that kept autointerp closed
The standard recipe for labelling features automatically has had two halves. An explainer model reads the text that activates a feature and writes a sentence describing it. A scorer model then reads that sentence and predicts the feature's activation on held-out text, token by token, and the correlation between predicted and real activations is the score. The second half, simulation scoring, is what made the recipe expensive, because scoring one feature means generating a predicted number for every token in every example.
Gonçalo Paulo, Alex Mallen, Caden Juang and Nora Belrose at EleutherAI put out a paper on October 17 with a pipeline that keeps the explainer half and replaces the scorer with cheaper tests. The whole thing runs on open models, with Llama 3 70B as the default explainer and scorer, and the code and explanations are public. The point of the paper is less any single result than that a group with a few GPUs can now label an entire dictionary.
Detection and fuzzing instead of simulation
The two new scorers are simple. Detection shows the scorer an explanation and a mix of sequences, five activating examples drawn across activation deciles and twenty random non-activating ones, and asks which sequences contain the feature. Fuzzing goes one step further by marking the activating tokens with delimiters and asking the scorer which examples are marked correctly, with two of every set deliberately mislabelled. Both reduce scoring to a classification task, which a model can do in one pass rather than one number per token.
The paper also adds intervention scoring, which asks whether steering with a feature changes the model's output in a way that matches the explanation, and a generation score, where the scorer writes text it thinks should activate the feature and you count whether it does. Intervention scoring is the one we find most interesting, because it evaluates what a feature does rather than what it reads, and the blog notes that some features promote or suppress output tokens without any interpretable input pattern at all. A simulation score would call those features uninterpretable. An intervention score can rescue them.
What the pipeline costs
The blog post gives per-million-feature prices that explain why this matters. Detection or fuzzing costs about 125 dollars with GPT-4o mini and 2,540 dollars with Claude 3.5 Sonnet. Generating the explanations costs 160 dollars and 3,400 dollars on the same two models. Simulation scoring costs 4,700 dollars and 96,000 dollars. For the roughly 1.5 million features in the GPT-2 dictionaries they cite, the full pipeline comes to around 1,300 dollars with Llama or 8,500 dollars with Claude, against something like 200,000 dollars for the earlier approach.
Those are API prices, and the open-model figure assumes you rent the compute, but the ratio is the point. A two-order-of-magnitude drop in scoring cost moves autointerp from a thing you do on a sample of features for a paper figure to a thing you do on every feature as a matter of course. That changes what an SAE release can include. A dictionary shipped with a label and a score for every latent is a different artefact from a dictionary shipped with weights.
How the scorer decides what is interpretable
The uncomfortable finding in the paper is that the scores do not agree with each other very well. Detection and fuzzing correlate with simulation at a Pearson coefficient of 0.61. That is high enough to say they measure something in common and low enough that the ranking of features by interpretability depends on which test you run. A feature with a crisp explanation that predicts presence but not magnitude will look good under detection and poor under simulation. A feature that fires broadly will do the reverse.
Two smaller results push the same way. Explanations written from a broad sample of activating examples across deciles beat explanations written from only the top activations, which suggests that the top-activation dashboards most of us use for a quick look are biased toward the sharpest and least typical behaviour. And chain-of-thought prompting for the explainer made explanations worse, with the blog saying the model tends to overthink and fix on extraneous details. Interpretability scores are measurements of a pair, the feature and the scorer, and reporting one without the other is not quite a result.
What we would want next
The obvious use of a cheap pipeline is to run it on everything and compare dictionaries by the distribution of scores, and the paper does this across SAE widths, activation functions and losses on two open models, finding that latents are much more interpretable than raw neurons. The result we want is the inverse. Take the features where detection, fuzzing and intervention disagree most, look at them by hand, and work out which scorer was right. If the disagreement is systematic, it tells us something about what SAEs are finding. If it is noise, it tells us how much to trust any single number.
We are going to run the pipeline on our own small dictionaries this quarter. The first question we will ask is how many of our features have a stable label across scorers, because that number, rather than the mean score, is the one we would put in a paper.
Sources
From the foundation