Gemma Scope puts sparse autoencoders in reach of anyone with a GPU
DeepMind released over 400 sparse autoencoders covering every layer of Gemma 2, trained with more than a fifth of the compute that went into GPT-3. What that changes for groups outside the big labs.
What was released
On 9 August DeepMind's interpretability team published Gemma Scope, a suite of JumpReLU sparse autoencoders trained on every layer and sub-layer of Gemma 2 2B and 9B, plus selected layers of the 27B model. The main release contains more than 400 autoencoders. Counting the variants at different sparsity levels there are over 2,000. Together they hold more than 30 million learned features, with the authors' own caveat that many of those features probably overlap.
The autoencoders come at three sites per layer: the residual stream, the MLP output, and the attention output. Widths run from 16,384 features up to a little over a million, in powers of two. The 2B model has 26 layers, the 9B has 42, and the 27B has 46, so the coverage for the two smaller models is close to complete. The weights are on Hugging Face under a CC-BY-4.0 licence, they load through the SAELens library, and Neuronpedia hosts an interactive browser for the 2B residual stream at layer 20.
Until this week, if you wanted to work with sparse autoencoders on a competent language model, you either trained your own or you read about what Anthropic had done to Claude. Now the artefact is downloadable. That is the whole point of the release and we think it is the most useful thing to happen in this corner of the field this year.
The cost that nobody outside a lab could pay
The numbers in the paper make clear why this was not going to happen from a university group. Training the suite used over 20 percent of the training compute of GPT-3. The team saved about 20 pebibytes of activations to disk along the way. Each autoencoder was trained on billions of tokens, four billion for the narrowest, eight billion for most, and sixteen billion for the million-width dictionaries.
For comparison, our own attempts at reproducing the 2023 monosemanticity results ran on a single card and a model with a few hundred neurons per layer. The gap between that and a full sweep of a 9B model at every layer is not something a grant closes. It is the kind of thing that only gets done once, by a group with idle accelerators, and then either gets shared or does not.
This is what makes the release consequential rather than merely generous. The autoencoder is the expensive part. Everything downstream of it, feature labelling, probing, steering, circuit finding, is cheap. By paying the expensive part once and publishing the result, DeepMind has moved the bottleneck for the whole community from compute to ideas.
JumpReLU and why the architecture choice matters
The autoencoders use a JumpReLU activation rather than the plain ReLU from last year. The paper gives the rationale as allowing greater separation between two jobs the encoder has to do: deciding which features are active, and estimating how strongly they are active. A plain ReLU forces those two decisions through the same threshold at zero. A JumpReLU has a learned threshold per feature below which the activation is clamped off, so a feature can be inactive at small positive pre-activations without distorting its magnitude when it is on.
Whether this is the right architecture is still open. There is a live argument in the field about gated and top-k variants, and about whether reconstruction loss is the objective we should care about at all. What the release does is settle the question of which architecture people will use for the next year, because the trained weights exist for this one and not for the others. That is worth being aware of. A tooling decision made for practical reasons can harden into a methodological default.
What we plan to do with it
The first thing is the boring thing. Load the layer-20 residual autoencoder for the 2B model, pick a few hundred features at random, and see what fraction we can label with confidence. The paper and the Neuronpedia demo show curated examples, and the curated examples always look good. The uncurated fraction is the number we want.
The second is a feature-consistency check across widths. The same layer has dictionaries at 16K, 32K, 65K, and up. If the method is finding real structure, a feature at 16K should split into a few related features at 65K, and the coarse feature should be close to a sum of the fine ones. If instead the dictionaries look unrelated, that is evidence the features are artefacts of the training run. This is a cheap experiment that nobody outside the lab could run before because nobody had a matched family of dictionaries.
The third is a probing comparison. Take a concept with a clean label, train a linear probe on the raw residual stream, and compare it with the best single feature and the best small set of features from the dictionary. We want to know whether the features buy anything over the probe on a task where we already know what we are looking for. We would not bet either way.
The ecosystem this sits in
It is worth naming the tools this release assumes. SAELens is the loading and training library the model card points to. Neuronpedia is the hosted feature browser, and it is where most people will first meet these dictionaries. DeepMind mentions an internal tool called Mishax for exposing activations. Around all of this sits the open-weights model itself, which is the real precondition. None of this works on a model you cannot run.
The pattern we expect to see is that Gemma 2 becomes the reference organism for interpretability the way certain strains become the reference organism in biology. Not because it is the best model but because it is the one with the most instrumentation. Papers will report results on Gemma Scope features because reviewers can check them, and results on proprietary models will be read as claims rather than findings.
The paper says the aim is to make more ambitious safety and interpretability research easier for the community. We think it will, and we also think the first few months of results will be mostly negative, because that is what happens when a method built and tested inside one group meets a hundred groups with different questions. That is fine. The negative results are the ones we most need and the ones that were hardest to produce when the only people who could run the experiments were the people who built the method.
Sources
From the foundation