Decomposing parameters instead of activations: stochastic parameter decomposition
Bushnaq, Braun and Sharkey propose SPD, which splits a network's weights into rank-one subcomponents chosen by stochastic ablation rather than gradient attribution. Method notes, why the authors think it sidesteps some SAE failure modes, and how far it has actually been tested.
A different object to decompose
Most of the interpretability tooling we use looks at activations. Sparse autoencoders take the residual stream at one layer and try to write it as a sparse sum of directions. Stochastic parameter decomposition, posted on June 27 by Lucius Bushnaq, Dan Braun and Lee Sharkey, looks at the weights instead. The goal is to write each weight matrix as a sum of rank-one subcomponents such that, for any given input, only a small number of them are needed to reproduce the network's output.
The authors frame this as the difference between explaining a program from its traces and explaining it from its source code. An activation decomposition tells you what representations exist at a point in the forward pass. A parameter decomposition tries to say which pieces of computation the network is made of and which of them fire on a given input. That is a stronger claim, and the interesting question is whether it can be made to hold on anything bigger than a toy.
How SPD picks its subcomponents
The predecessor method, attribution-based parameter decomposition or APD, used gradient attributions to estimate how important each subcomponent was on each input. SPD replaces that with stochastic masking. For each input, a small MLP predicts an importance score for every subcomponent. The training loop then samples random ablation masks in which subcomponents predicted to be unimportant are masked out more aggressively, runs the masked network, and penalises any change in output.
The effect is that importance is measured by what happens when you actually remove things, averaged over many random ablation patterns, instead of by a first-order gradient estimate. The authors say this gives a more direct causal measurement and removes a failure of APD where subcomponents shrank towards zero during training. They also report that SPD is less sensitive to hyperparameters and can decompose multi-layer networks where APD could not.
Why the authors think this addresses SAE failures
Two of the best-known problems with sparse autoencoders are feature splitting and feature absorption. Feature splitting is what happens when you grow the dictionary and a single feature fractures into several finer ones with no principled stopping point. Absorption is when a general feature stops firing on inputs that also trigger a more specific feature, so the general concept appears to have holes in it. Both are symptoms of the same thing. An SAE's decomposition depends on the dictionary size you chose, and there is no ground truth telling you which size is right.
The SPD pitch is that parameters give you a different kind of ground truth. The network's weights are finite and fixed, and the objective asks for the most parsimonious set of rank-one pieces that reproduce its behaviour across inputs. The authors argue this yields a decomposition that explains feature geometry and computational structure rather than a resolution-dependent picture of the activations. We find the argument plausible, and we also note that it is an argument rather than a measurement. Nobody has yet shown that SPD avoids splitting and absorption on a real language model, because nobody has run it on one.
What it has been tested on
The validation is entirely on toy models with known ground truth. On the toy model of superposition, SPD recovers the individual feature directions. On a model whose target is an identity matrix, it recovers the expected components. On a three-layer network it decomposes mechanisms that span layers, which is the case APD failed on. These are the right first tests, because you cannot tell whether a decomposition method is correct unless you already know the answer.
The authors are clear about the limits. SPD produces rank-one subcomponents, and grouping those into mechanisms a human would recognise requires a post-hoc clustering step that is not part of the method. The compute cost is not trivial, since every training step runs multiple masked forward passes. Hyperparameter sensitivity is reduced relative to APD but still present. And the largest network in the paper is a toy. The post says the team is working on the modifications needed for language models, with arithmetic, syntax and factual recall as the target capabilities to decompose.
What would convince us
The test we want is a small transformer, on the order of a few million parameters, trained on a task where we already have a reasonably trusted circuit-level account from other methods. Run SPD, cluster the subcomponents, and check whether the clusters line up with the known circuit. Then deliberately vary the number of subcomponents and the masking schedule and see whether the recovered mechanisms stay stable. Stability under those choices is the property SAEs lack, and it is the property SPD is claiming.
If that holds, the second test is the absorption case directly. Build a network with a known general feature and a known specific one that co-occur, and check whether SPD assigns them to separate subcomponents that both fire on the overlapping inputs. A clean result there would be the first real evidence that decomposing weights buys something activations cannot. Until then we are treating SPD as a promising method with a strong motivating argument and a toy-scale track record.
Sources
From the foundation