Crosscoders and model diffing: features that span layers and checkpoints
Anthropic trained one sparse dictionary across every layer of a model, and then across a base model and its finetuned version. Notes on what a crosscoder is, what it found, and why comparing a model against its predecessor may be the most useful thing dictionary learning does.
One dictionary for the whole residual stream
A sparse autoencoder reads activations at one layer and writes a reconstruction of that same layer. A transcoder reads one layer and writes the next. The research update Anthropic published on October 25 describes a third variant, the crosscoder, which reads from many layers at once and writes a reconstruction of every one of them. A single feature therefore has an encoder vector per layer and a decoder vector per layer, and the sparsity penalty is applied to the feature, not to any one layer's copy of it.
The authors are careful to label this a preliminary note, presented in the spirit of a lab meeting rather than a full paper. We want to hold on to that framing, since the evidence for the claims here is early.
The motivation is cross-layer superposition. If a model has more layers than a given circuit needs, it can spread one computation across two or three adjacent layers that behave almost like parallel branches, since the residual stream is linear and each layer just adds to it. A per-layer SAE would then see fragments of a feature at each layer and learn several partial versions of it. A crosscoder can learn it once.
How the loss is set up
The reconstruction term is the sum of per-layer mean squared errors. The sparsity term weights each feature's activation by the sum over layers of its decoder norms, an L1 of norms rather than an L2. The authors give two reasons for that choice. The L1 version makes the crosscoder's loss directly comparable to the summed losses of per-layer SAEs trained with the same penalty, and it gives no explicit incentive to smear a feature across layers, which the L2 version does.
That second property matters later. In the diffing experiments they report that the L1-of-norms version surfaces a mix of shared and model-specific features, while the L2-of-norms version tends to find only shared ones. The L2 version does optimise the MSE against global L0 frontier more efficiently, so the choice depends on whether you are after efficiency or after layer-specific and model-specific structure.
There are also three causality variants. An acausal crosscoder reads all layers and writes all layers. A strictly causal one reads earlier layers and writes later ones, which makes it a generalised transcoder. A weakly causal one sits in between. Most of the reported experiments use the acausal version.
What the 18-layer experiment showed
The first experiment trains a global acausal crosscoder on the residual stream at every layer of an 18-layer model, and compares it against 18 SAEs trained separately, one per layer, with the same L1 coefficient and per-layer activation normalisation. Controlling for the total number of features across layers, the crosscoder reaches substantially lower eval loss than the per-layer SAEs. The authors read this as evidence of linearly correlated structure across layers that the crosscoder captures as single cross-layer features.
The cost side is less flattering. Measured in training FLOPs, the crosscoder is less efficient than the per-layer SAEs at reaching the same loss, by a factor of about two at large compute budgets. So the method wins on features per unit of reconstruction and loses on compute per unit of reconstruction.
Looking at decoder norms across layers for 50 randomly sampled features, most peak at a particular layer and decay on either side. Sometimes the decay is sharp, which is a localised feature. Often it is gradual, with many features carrying substantial norm across most or even all layers. That is the picture the cross-layer superposition story predicts, though the note does not put a number on what fraction of features fall into each group.
Diffing Claude 3 Sonnet against its base model
The experiment we find most useful is the smallest one. The authors trained a crosscoder with one million features on the middle-layer residual stream of Claude 3 Sonnet and of the base model it was finetuned from. The question was whether the dictionary would split cleanly into features shared by both models and features specific to one of them. Looking at the relative decoder norms in the two models, the features fell into three obvious clusters, base-specific, finetuned-specific, and shared.
The model-specific groups are small, between four and five thousand features for each model out of a million. Among the finetuned-only features they picked out a refusal feature that fires on dangerous requests, a code review feature that responds to requests for feedback on code, and a feature that fires when a human asks the Assistant personal questions about itself. Among the base-only features there is one for language models being cast as a character in a system prompt, and one for dialogues with a smartphone assistant. The authors say plainly that these are cherry-picked and that most model-exclusive features are not immediately interpretable.
For the shared features they checked whether the decoder directions in the two models line up. Almost all are highly aligned, which supports reading them as the same concept doing the same job in both models. A few thousand have very low or negative correlation, which the authors have not investigated and tentatively read as concepts the finetuned model reuses in a new way.
Why diffing may be the practical use
A full feature dictionary for a frontier model is a large object with millions of entries, most of which nobody will ever look at. A diff is a small object. Five thousand features that changed under finetuning is a list a team can actually read, and it is a list that answers a question people already ask, namely what did this training run do to the model. The note points to comparing finetuning strategies as the immediate applied case, and to RL finetuning as the longer-term safety case, since many arguments about risk single out models that were finetuned with RL.
There is a competing approach worth keeping in view. The note cites work by Kissane et al. showing that SAEs trained on a base model often transfer well to the finetuned model, so one can diff by asking which base-model features are active in one model but not the other on a given prompt. The crosscoder describes the diff instead in terms of model-specific abstractions that the base dictionary would never contain. Both views are useful and they answer different questions.
What we would want to see next is a diff between two checkpoints in the middle of a training run, with a known intervention between them, to check that the model-specific cluster contains what the intervention should have added. Until someone runs that control, the three-cluster picture is suggestive rather than validated.
Sources
From the foundation