The problem in one sentence

Chris Olah posted a note on the Transformer Circuits Thread on August 7 with a claim that is easy to state and uncomfortable to sit with. A sparse replacement for a layer can be monosemantic, can reach low reconstruction error on the training distribution, and can still compute its outputs through a different mechanism than the layer it replaces. Sparse autoencoders and transcoders are judged almost entirely on reconstruction error and interpretability, and neither of those checks the mechanism.

The reason this matters for us is that everything downstream of a transcoder inherits the assumption. When we trace an attribution graph through a replacement model, we are reading off the mechanism of the replacement, and we are trusting that it is the mechanism of the original. If the two differ in ways that the on-distribution error does not reveal, the graph is a faithful description of the wrong object.

The setup

The toy model is as small as it could be. A one-layer network computes the absolute value of each input coordinate using two ReLU neurons per coordinate, so that y equals ReLU of x plus ReLU of minus x. Inputs are drawn uniformly from minus one to one with a controlled sparsity, so most coordinates are zero most of the time. This is a model whose true mechanism we know completely, which is the whole point. There is no ambiguity about what a faithful transcoder should learn.

Then a special datapoint is injected. Call it p, a fixed vector with a few nonzero coordinates, and make it appear in the training data with some fraction of the time. The transcoder is trained to read the layer input and predict the layer output, with a tanh L1 penalty that Olah reports behaves closer to an L0 penalty than plain L1 does. He notes that plain L1 produced messier duplicate features and excess capacity, so the choice is not cosmetic.

The datapoint feature

When the repeat fraction and the sparsity coefficient are both high enough, the transcoder grows a feature that fires only on p and reproduces p. The form is y plus p times ReLU of lambda p dot x minus b. It is a memorisation feature. It is monosemantic, it lowers reconstruction error, and it is completely unfaithful, because the underlying model has no notion of p at all. The true model computes absolute value coordinate by coordinate and treats p like any other input.

The note's hyperparameter scans suggest the two factors multiply rather than acting as independent thresholds. More repetition makes the feature more likely, stronger sparsity pressure makes memorisation cheaper relative to reconstructing p through the genuine absolute value features, and the product controls whether the datapoint feature appears. We find that framing clarifying. It says the failure is a property of the training objective and the data together rather than a quirk of one setting.

Why off-distribution is where it bites

On the training distribution, the memorising transcoder and the faithful one produce nearly the same outputs, which is why the usual metrics do not separate them. The difference shows up on inputs near p but not equal to it. The faithful mechanism handles those by the same absolute value computation it uses everywhere. The memorising feature either fires when it should not or hands the input back to the remaining features which were never fully trained to cover that region. Olah's phrasing is that alternate mechanisms may generalise differently off distribution, and that is the entire threat.

This is what we meant about the quiet assumption. The strongest hope for mechanistic interpretability is an explanation reliable enough to predict behaviour outside the data you trained the explanation on. A replacement model that is accurate on distribution and mechanistically different is exactly the kind of explanation that fails that test while passing every metric we normally report. Real language models have a great deal of repeated data, and real dictionaries are trained under strong sparsity pressure, so the two ingredients are present at scale.

What we would try

The note floats Jacobian matching as a possible fix, meaning train the replacement so that its local gradient structure matches the original's rather than only its outputs. A memorising feature has a very different Jacobian around p than an absolute value computation does, so the penalty would see what the reconstruction loss cannot. Olah presents this as preliminary and we would treat it that way, but it is the natural experiment and the toy model makes it cheap.

The bigger open question is prevalence. Nothing in the note tells you how often datapoint features, or their more subtle relatives, appear in transcoders trained on a real model. We would start by looking for features in existing dictionaries whose activation set is a handful of near-duplicate training strings, and then check whether the attribution graphs that pass through them still hold up when the string is perturbed. If the graphs survive, the concern is theoretical. If they do not, we have been reading some of our circuits off the wrong model.

Sources

  1. Chris Olah, Mechanistic (Un)Faithfulness in SAEs and Transcoders: A Toy Model (Transformer Circuits Thread)