Transcoders: making MLPs legible without the activations
Dunefsky, Chlenski and Nanda train sparse replacements for MLP layers that map inputs to outputs directly. The payoff is that feature-to-feature connections can be read from the weights, before you ever look at an example.
The problem with reading through an MLP
Sparse autoencoders have given us features in the residual stream that are far more interpretable than neurons. They have not given us circuits. The reason is the MLP layer. A feature that is a clean direction in the residual stream before an MLP gets smeared across a very large number of neurons inside it, and to explain why a feature after the MLP fired you have to reason about all of those neurons and their nonlinearities. An SAE sees the activations on either side. It says nothing about the function in between.
Jacob Dunefsky, Philippe Chlenski and Neel Nanda posted a paper this week that attacks this directly. Instead of training a sparse model to reconstruct a layer's activations from themselves, they train a sparse model to reconstruct the MLP's output from the MLP's input. They call it a transcoder. It is a wide, sparsely activating one-hidden-layer ReLU MLP with an encoder and a decoder, trained on the usual combination of reconstruction loss and an L1 penalty on the hidden activations.
What changes when features map inputs to outputs
The trick that makes transcoders more than a reformulated SAE is in how attribution factorises. Suppose transcoder feature j in a later layer fired, and you want to know how much an earlier transcoder feature i contributed. The contribution is the product of two terms. The first is feature i's activation on this input, which is input-dependent. The second is the dot product of feature i's decoder vector with feature j's encoder vector, which depends only on the weights and is the same for every input.
That second term is what the authors call input-invariant, and it is the thing an SAE-based circuit analysis cannot give you. With SAEs, every claim about how features connect has to be established by running examples and patching. With transcoders, you can take feature j, look at the encoder-decoder dot products with every feature in an earlier layer, and read off which earlier features tend to drive it, before running a single example. The examples then tell you which of those connections were live for a particular prompt.
We want to be careful about what this does and does not cover. The factorisation handles the path through MLPs. Attention is treated as fixed: the analysis uses the actual attention patterns from the forward pass and does not explain why the model attended where it did. The authors say so plainly.
Do you lose anything by replacing the MLP
The obvious worry is that a transcoder is a worse approximation than an SAE, because it has to learn a function rather than an identity. The paper trains transcoders on GPT-2 small, Pythia-410M and Pythia-1.4B, with sweeps over the L1 coefficient to trace out the sparsity-fidelity frontier, and compares to SAEs trained on the same layers. The transcoder frontiers match or dominate the SAE frontiers, and the gap favours transcoders more on the larger models.
On interpretability, they had humans rate a sample of features blind. Of 50 transcoder features, 41 were judged interpretable, against 38 of 50 SAE features. That is a small sample and we would not read much into the three-feature difference, but it rules out the failure mode where transcoders are faithful and meaningless. On cross-entropy loss recovered when the transcoder is spliced into the model in place of the MLP, there was no meaningful difference from SAEs.
The greater-than circuit, again
The case study is the greater-than task in GPT-2 small: given a sentence like the war lasted from the year 1732 to the year 17, predict a two-digit continuation greater than 32. The original circuit analysis of this task found it spread across many MLP neurons. With transcoders, the authors find that about 24 features in the layer 10 transcoder are enough to recover most of the task performance. The features fire on specific ranges of the starting year and boost the logits of years above that range.
The point of the case study is less the circuit itself, which was already known, and more the size of the description. Far fewer transcoder features than neurons are needed to account for the behaviour, and the input-invariant weights tell you which of them push in which direction without a per-example patching campaign.
What we would want to know next
The honest limitations are the ones the authors list. Transcoders are approximations, and whatever they fail to reconstruct is invisible to any circuit you build from them. The case studies are qualitative and few. And the method does not extend to attention, because a softmax over query-key products is not something a sparse ReLU layer can replace in the same way.
The part of this we expect to matter most is the idea of replacing a component of the model with an interpretable stand-in and then analysing the replacement. If you can do it for one MLP, you can in principle do it for every MLP at once and have a replacement model whose feature connections are all readable from weights, with the original attention patterns supplying the rest. Whether the error terms from stacking many approximations stay small enough for that to be useful is the experiment we would run first.
Sources
From the foundation