A latent that looks right and is not

David Chanin, James Wilken-Smith, Tomas Dulka, Hardik Bhatnagar, Satvik Golechha and Joseph Bloom posted a paper this month with an example that we have not been able to stop thinking about. In the Gemma Scope layer 3 residual stream SAE with 16k latents and L0 of 59, latent 6510 looks like a detector for tokens that start with the letter S. On a first letter classification task it reaches an F1 of 0.81. Then you feed it the token short and it does nothing.

A different latent, 1085, fires instead. That latent is aligned with variants of the word short. Its cosine alignment with the starts-with-S probe direction is only a fifth of the main latent's, and yet on that token it activates 55 times more strongly. When the authors ablate the probe-aligned component of latent 1085, the model's spelling knowledge for that token goes with it. The starts-with-S direction is really there in the activation. It has simply been folded into a token-specific latent, and the latent labelled starts-with-S has learned to stay quiet.

Why sparsity makes this happen

The paper's explanation is short enough to reproduce here. Take two features in a hierarchy, where feature f1 only ever occurs together with a parent feature f0, and both fire with magnitude one. Without the hierarchy an SAE recovers both as separate latents. With the hierarchy, the SAE can do something cheaper. It lets the child latent absorb the parent direction into its decoder, and it trains the parent latent's encoder to fire only when the parent is present without the child. Now hierarchical inputs light up one latent instead of two.

Reconstruction loss is unchanged by this move, because the parent direction is still being written out, just by a different latent. Sparsity loss falls, and the authors show in an appendix that its derivative with respect to the degree of absorption is minus the co-occurrence probability. Whenever the two features co-occur at all, gradient descent prefers more absorption. The objective asks for it directly.

How they measured it

The test bed is first letter identification. Prompts of the form token has the first letter, followed by the capital, give a task where the ground truth is unambiguous and a linear probe on the residual stream does well. The authors run k-sparse probing to find the latents that split a letter across several latents, then look for tokens where the probe succeeds but every main latent fails. On those tokens they run integrated gradients ablation. If some other latent has a larger ablation effect than the main ones and a cosine similarity above 0.025 with the probe direction, that counts as an absorption. The absorption rate is that count divided by the probe's true positives.

They call this a conservative baseline, and they are right to. It misses cases where several latents share the absorbed direction, and cases where the main latent fires weakly rather than not at all. Even so, absorption showed up in every language model SAE they tested. The primary results are on Gemma-2-2B with Gemma Scope SAEs at 16k and 65k width, with confirmation on Qwen2 0.5B using their own L1 SAEs and on Llama 3.2 1B with both L1 and TopK architectures. Past layer 17 of Gemma-2 the metric becomes unreliable, because attention has moved the first letter information elsewhere by then.

Width and sparsity do not rescue you

The natural hope is that this is a small SAE problem and a bigger dictionary will pull the parent feature back out. The results say otherwise. Wider and sparser SAEs show higher rates of absorption, and the rate rises with L0 sparsity in the later layers where the sparsity pressure is strongest. On the underlying classification task no SAE latent matched the linear probe. Low L0 SAEs learned high precision, low recall latents for letters, and high L0 SAEs learned the opposite, with the best F1 sitting around L0 of 25 to 50 in early layers and 50 to 100 later on.

That last point deserves attention on its own. A single linear direction in the residual stream carries the first letter cleanly enough for a probe to read it. The SAE, which is supposed to give us that direction as an interpretable atom, gives us instead a family of latents that between them cover it, none of which is trustworthy alone. Interpretability that depends on reading one latent at a time inherits every gap in that coverage.

What this does to the dictionary story

The appeal of dictionary learning has been that a latent with a clean activation pattern and a clean label can be treated as a variable in the model's program. Absorption breaks the step from clean looking to clean. Latent 6510 has a good F1, sensible top activations and an obvious name, and it will still be silent on an input that plainly belongs to its category. If you built a monitor on it, the monitor would have holes shaped like whichever tokens happened to earn their own latent. If you steered with it, you would miss those same tokens.

The paper does not have a fix, and it says so. Varying size and sparsity is not enough, and the authors suggest that architectures which represent hierarchy explicitly, or objectives that do not reward merging a parent into its children, are the places to look. Our own reading is that any SAE evaluation which reports only reconstruction and interpretability scores is now incomplete. An absorption rate, on a task with a known ground truth direction, belongs next to them. Until it is standard, we would treat a feature's label as a hypothesis about where it fires, and go looking for the tokens where it should fire and does not.

Sources

  1. Chanin et al., A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders (arXiv 2409.14507)