Model diffing with crosscoders: why the features unique to one model are the hardest to read
Anthropic's crosscoder diffing update compares two models through one shared dictionary. The features that belong to only one model come out dense and polysemantic, and the proposed fix says something about how dictionaries allocate capacity in general.
The setup
A crosscoder is a sparse autoencoder trained on the activations of more than one model at once. You feed it the residual stream from model A and model B at the same token position, and you train a single dictionary whose features have separate decoder directions for each model. The loss is the usual reconstruction plus sparsity, summed over the two models. The point is that if a feature exists in both models it gets one dictionary entry with two decoder vectors, and if it exists in only one model the decoder vector for the other model should go to zero.
That gives you a diff for free. Compare the decoder norms across models and you can sort every feature into shared, A-exclusive or B-exclusive. If A is a base model and B is a fine-tune of it, the B-exclusive features are your candidate list for what fine-tuning added. This is the pitch from Anthropic's earlier crosscoder work, and the February update from Siddharth Mishra-Sharma, Trenton Bricken, Jack Lindsey and colleagues is a progress report on what happens when you actually try to read those exclusive features.
What the exclusive features looked like
They were hard to read. The team reports that features exclusive to one model tend to be more polysemantic and more dense in their activations than shared features, which makes them difficult to interpret. That is the opposite of what the method promises. The whole reason to do model diffing is to get a short list of clean, nameable differences, and the list you get instead is a set of features that fire often and on many unrelated things.
This is a preliminary result and the authors frame it as such. The update is written in the spirit of sharing partial findings with the field rather than a finished paper. But the pattern was clear enough that they went looking for a cause, and the cause turns out to be structural rather than a matter of any one model pair.
Why it happens: capacity competition
The team built toy models to isolate the mechanism, and the explanation is about competition for a fixed feature budget. A dictionary has a limited number of features. A shared feature earns its place by explaining activation patterns in both models, so it pays off twice against the same sparsity cost. An exclusive feature only explains one model, so to justify its slot it has to carry more information. In practice that means it absorbs several unrelated directions and fires on all of them. Polysemanticity is the symptom, and the cause is that the training objective quietly prices exclusive features at a discount.
We find this a useful way to think about what any sparse dictionary is doing. We tend to talk about SAE features as if they were discovered, and the discussion is about whether the discovery is real. This result reminds you that the features are the solution to an optimisation problem under a budget, and when you change the problem, for instance by asking one dictionary to serve two models, the budget gets spent differently. Nothing about the underlying model changed. The allocation did.
The proposed fix
The intervention is simple to state. Designate a subset of features as shared in advance and give them a lighter sparsity penalty. Because the shared features are now cheap, they soak up everything common to both models without needing to compete, and the remaining features are free to specialise on what actually differs. The update reports that with this change the exclusive features became more monosemantic and interpretable, and the team could isolate meaningful behavioural differences between the two models.
Notice what the fix does and does not do. It does not add capacity. It changes the relative price of two kinds of features so that the optimiser stops overloading the expensive kind. That is a general lever, and we would expect versions of it to matter anywhere a dictionary is asked to cover heterogeneous data, including single model SAEs trained across very different data distributions or across layers.
What we would want to see next
The obvious question is how sensitive the result is to the number of designated shared features and the size of the penalty discount. If there is a wide plateau where exclusive features stay clean, this is a dependable trick. If there is a narrow band, it is a hyperparameter you tune per model pair and the method is much less turnkey than it looks. The update does not settle this and we would treat it as the first thing to check in a replication.
The second question is whether the exclusive features found after the fix are the same ones a different approach would find. Training a plain SAE on the fine-tuned model and looking for features absent from the base model's SAE is a cruder diff, but it does not suffer from this particular capacity competition. If the two methods agree on the top differences, that is good evidence the fix is surfacing something real about the fine-tune rather than something about the crosscoder. That comparison is cheap and we have not seen it published yet.
Sources
From the foundation