SAEs can steer after all, if you pick the feature properly
A rebuttal from DTU revisits the AxBench finding that sparse autoencoders lose to simple steering baselines. With supervised feature selection and labelling, SAE steering on Gemma 2 lands close to LoRA. Notes on how much of a negative result was really about feature selection.
The result being contested
When AxBench came out in early 2025, the line everyone repeated was that even simple baselines outperform sparse autoencoders for steering. On that benchmark a model is prompted with Alpaca-Eval instructions, steered toward a concept, and its outputs are scored by GPT-4o on fluency, instruction following and whether the concept actually shows up, combined by a harmonic mean. Prompting won easily. LoRA and ReFT did well. SAE feature steering sat near the bottom, and a good number of labs quietly deprioritised SAE work partly on the strength of that chart.
A paper posted on May 29 by Mikkel Godsk Jørgensen and Lars Kai Hansen at DTU Compute argues that the SAE row in that chart was mostly measuring the quality of the feature labels. Their claim is narrow and they state it plainly. With a different way of choosing and naming features, SAE steering on the same benchmark comes close to the LoRA reference. It still does not match prompting, and they say that too.
Labels from data instead of from an LLM
The Neuronpedia labels used in the original AxBench comparison come from a language model reading top activating examples and writing a description. Jørgensen and Hansen replace that with a labelling pipeline built on a multi label dataset. They use Stack Exchange dumps processed with EleutherAI's tooling, where each post carries user assigned tags such as baseball or mlb, and they pick forums on academia, biology, chemistry, cooking, computer science, history, law, literature, physics, politics and sports.
Each SAE feature is treated as a linear probe with no known label. For every feature and every tag they compute how often the feature fires on posts with that tag, threshold it, and score the match with a calibrated F1 that corrects for how common the tag is. Then they denoise. Tags with fewer than 50 supporting posts are dropped, features are filtered by the output score from Arad and colleagues as a proxy for steerability, and the top 50 feature label pairs are kept. The support threshold and the calibration parameter were tuned on layer 32 only, and layer 17 was held out to check the method transfers.
What the steering runs showed
The model is Gemma-2-9b-it with Gemma Scope SAEs of width 131k, and the steering method is the clamp from Templeton and colleagues, with the feature activation set to a value alpha and the reconstruction error added back. Rather than fixing one alpha for all features, as AxBench did, they sweep alpha from 100 to 575 per feature and pick with repeated cross validation, 10 repeats of 4 splits, reporting on the held out partition. Each feature label pair is steered on 20 random instructions.
The aggregated ratings for SAE steering with the selected features land near the LoRA and ReFT reference values on several forums at both layers, and far above the reproduced Neuronpedia random baseline, which sits close to where AxBench reported SAEs. Prompt steering also improves over the AxBench numbers, which the authors attribute to simpler labels being easier for the subject model to follow. To check that the improvement is not only the judge liking the new labels, they look at the concept rating alone, the one component that depends on the label, and find the gap between feature steering and prompt steering mostly closed there too.
The ablations are where it gets interesting. Using high sparsity SAEs, with low L0, did not much change the steering result, which cuts against the earlier Wang and colleagues finding that low L0 features steer better. And when they compared randomly chosen features to pipeline chosen features while fetching the original Neuronpedia labels for both, the pipeline picks still steered better. Selection is doing real work independent of the new names.
How much of the negative result was selection
Our reading is that AxBench established something true and something incidental at the same time. The true part is that SAE steering, done by taking a public feature dictionary and its auto generated labels, underperforms methods with trainable parameters. The incidental part is that this measured a specific labelling pipeline, and a pipeline with access to tagged text finds features whose causal effect matches their name much more reliably.
There are limits worth stating. One model, one benchmark, one judge. The authors ran a small trial with Danish undergraduates to check for judge bias against particular methods and found no clear evidence, but they call the sample small. They also observed cases where AxBench gave a concept score of zero to outputs on a semantically adjacent or even matching topic, which is a benchmark noise problem that affects every method. And the whole thing depends on Stack Exchange overlapping enough with whatever Gemma Scope was trained on, which is not public.
What we would do with this
Nobody should ship SAE steering in production on this evidence. The authors describe feature steering as controlled damage and they are right. What the paper restores is the research use. If a feature selected by tag data steers the model toward that tag, then the label is more than a story about top activations, and that is the property interpretability researchers actually need from a dictionary.
The experiment we want next is to run this pipeline against the same layer with a different labelled corpus, say Wikipedia categories rather than Stack Exchange tags, and see whether the same features get the same names. If they do, the negative result from 2025 was mostly about selection. If they do not, it was about the dictionaries, and we are back where we were.
Sources
From the foundation