What shrinkage is

A sparse autoencoder is trained on two terms, a reconstruction loss and an L1 penalty on the feature activations. The penalty is what makes the dictionary sparse. It is also a constant pressure on every active feature to be a bit smaller than it should be, because shrinking an activation always reduces the penalty and only sometimes hurts reconstruction. The result is that a standard SAE systematically underestimates feature magnitudes. The paper calls this shrinkage and it is a property of the objective, not of any particular training run.

The authors measure it with a quantity they call relative reconstruction bias, the ratio of the expected squared norm of the reconstruction to the expected dot product of the reconstruction with the true activation. An unbiased reconstruction gives one. A value below one means the SAE is reconstructing a scaled down version of the input. Baseline SAEs come out below one. The gated version comes out at about one.

The gated architecture

The fix separates two jobs that a standard encoder does with a single ReLU. One job is deciding which features are active. The other is estimating how active they are. The gated encoder has a gate path, which is a linear map followed by a step function that outputs zero or one, and a magnitude path, which is a linear map followed by a ReLU. The feature activation is the product. The L1 penalty is applied only to the gate path, so the sparsity pressure never touches the magnitude estimate.

Doubling the encoder would double the parameter count, so the authors tie the weights. The magnitude weights are the gate weights scaled elementwise by the exponential of a learned vector, which adds only two parameters per dictionary feature. The paper notes that with this tying the encoder is mathematically equivalent to a linear map followed by a jump ReLU, an activation that is zero below a learned threshold and identity above it. That equivalence is a hint about where the field goes next, since a threshold activation can be trained in more than one way.

There is one more piece. Because the step function has no gradient, the gate path is trained through an auxiliary loss that reconstructs the input using the pre-activation of the gate and a frozen copy of the decoder. The frozen copy is there so that the sparsity penalty on the gate cannot bias the decoder. This costs an extra decoder pass, about 50 percent more compute for the SAE itself, though the authors point out that most wall clock time in practice goes to generating the activations in the first place.

What the results show

The evaluation runs on three models. GELU-1L is a one layer model used for comparison with prior work. Pythia-2.8B is evaluated at five layers across MLP, attention and residual stream sites. Gemma-7B is evaluated at four layers across the same kinds of sites. At every site the gated SAE recovers more of the model's loss at a given sparsity, and the abstract's summary is that in some cases the gated SAE needs about half as many firing features to reach the same reconstruction fidelity. The paper describes this as a Pareto improvement over the prevailing training method within typical hyperparameter ranges.

Interpretability was checked with humans. Five raters looked at 150 Pythia-2.8B features and seven raters looked at 192 Gemma-7B features in a paired design, labelling each as interpretable, not interpretable or unsure. The one sided test came out at p equals 0.060 with a confidence interval on the mean difference of zero to 0.26. The authors say they cannot claim gated features are more interpretable and that they are at least comparable. We appreciate that they published the number rather than the conclusion.

What the authors say is still wrong

The limitations section is short and specific. The architecture is more complicated to train than a plain SAE. The jump ReLU is discontinuous, which is a problem for attribution methods like integrated gradients that are common in circuit analysis. The whole approach still rests on the sparsity and linearity assumptions about model computation that motivate SAEs in the first place. And the evaluations do not test whether the learned features are causally meaningful intermediate variables in the model's computation, which is the thing anyone using them for circuit discovery actually wants.

That last point is the one we would underline. Loss recovered and L0 are proxies. They are the proxies everyone uses, and a Pareto improvement on them is a real advance, but a dictionary can reconstruct well with features that do not correspond to anything the model uses. Nothing here tells you which case you are in.

Why architecture became a field

Until this paper the working assumption was that a sparse autoencoder was a fixed recipe and the interesting questions were about what it found. This paper shows that a specific defect in the recipe, shrinkage, is measurable, has a cause, and can be fixed with a small change that improves the headline metrics. Once that is established the recipe stops being fixed. Every term in the objective and every nonlinearity in the encoder becomes a knob, and the sparsity versus fidelity curve becomes a leaderboard.

The equivalence to a jump ReLU is the thread we would pull. If a thresholded linear encoder is the thing that works, the question becomes how to train the threshold directly, and whether the gate and magnitude decomposition is a convenient training trick or the actual mechanism. We would also want to see the same experiment repeated at much larger dictionary sizes, since capacity competition between features changes what shrinkage costs. And we would like to see anyone using these SAEs for circuit work report whether the gated features behave better under intervention, because that is the evaluation the paper says it did not run.

Sources

  1. arXiv: Improving Dictionary Learning with Gated Sparse Autoencoders (Rajamanoharan et al.)