The bet

Most of the interpretability field has spent the last two years trying to decode dense models after the fact. Train a sparse autoencoder on the activations, hope the dictionary it learns lines up with the concepts the model uses, and then try to trace circuits through those features. The paper Leo Gao and five coauthors posted on November 17 makes a different bet. Instead of decoding a dense model, train a model that is sparse from the start, so that each neuron has only a handful of connections and the circuits are there to be read off rather than reconstructed.

The mechanism is a constraint on the weights. At the strongest setting roughly one in a thousand weights is nonzero. That weight sparsity induces activation sparsity on its own, with about one in four activations nonzero, and the paper tracks this by watching the kurtosis of activations rise as the number of nonzero weights falls. The models are small by any current standard, with a typical configuration of eight layers, a 2048-wide hidden dimension and a 256-token context, and the evaluation tasks are all Python code completion.

What the circuits look like

The evaluation is a set of 20 hand-built Python completion tasks, and the ones the paper walks through are the kind you could explain to a first-year student. Closing a string with the right kind of quote, single or double. Predicting whether the next token is one closing bracket or two. Tracking whether a variable holds a set or a string. Telling if from while from for. Distinguishing a lambda from a def.

For each task the authors prune the sparse model down to the smallest subgraph that still does the job, and the subgraphs are tiny. The quote-closing circuit has 12 nodes and 9 edges and completes the task almost perfectly. The bracket-counting circuit runs through an averaging mechanism over earlier tokens. The variable-type circuit is a two-hop copy implemented by two attention heads with about 100 edges in total. Against dense models trained to the same pretraining loss, pruning the sparse model gives circuits roughly 16 times smaller. And the nodes have names you can defend. The paper describes residual channels that act as a quote-type classifier, which is the kind of claim that is easy to make and rare to be able to check.

The check is where we would push. The circuits are validated with mean ablation, which is a reasonable proxy for faithfulness and not the same as showing that the rest of the model is doing nothing on the task. The authors say as much and name causal scrubbing as the stronger test they did not run. They also note that features are not fully monosemantic, that some concepts smear across several nodes, and that the magnitude of a feature carries information beyond whether it is on. So even in a model built to be legible, the legibility is partial.

What it costs

Everything in the paper is bought with capability and compute. Holding model size fixed and increasing sparsity makes the model worse and its circuits smaller, so there is a straightforward tradeoff curve. Growing the model at a fixed number of nonzero weights improves both at once, which the authors attribute to the extra freedom in where the nonzero weights go. That second finding is the encouraging one, and it is also the one that comes with the largest asterisk.

The asterisk is compute. Unstructured sparsity defeats the tensor cores that make dense matrix multiplication cheap, and the paper reports that its sparse models need 100 to 1000 times more training and inference compute than dense models of equivalent loss. That is the difference between a research artefact and something you could deploy. The authors suggest semi-structured 2:4 sparsity or a sparse mixture of experts as ways to claw some of it back, but neither has been tried here.

Bridges to the models people actually use

The part of the paper that makes it more than a curiosity is the bridge. The idea is to train a sparse model alongside a dense one and learn an encoder and decoder at each sublayer that map between their activations, with a loss that combines normalised mean squared error on the activations and a KL term through hybrid forward passes. If the map is good, you can find a feature in the sparse model and use it to perturb the dense one.

The result is preliminary and stated as such. On two test tasks, manipulating a sparse-model channel such as the quote-type classifier steers the dense model's output in the expected direction. That is the first evidence we have seen that a sparse-by-construction model could serve as an interpretable proxy for a dense one rather than a replacement for it, and it is the direction we would want the follow-up work to take, because it sidesteps the compute problem by keeping the deployed model dense.

The open question

The authors are direct about the limit. Scaling sparse models beyond tens of millions of nonzero parameters while keeping them interpretable remains a challenge, and they expect that complex behaviours will produce circuits too large to read even with sparse training. The 20 tasks are simple by design. Nobody has shown that a sparse model can do in-context learning or multi-step reasoning with circuits a human can hold in their head, and there is a reasonable argument that polysemanticity is what efficient computation looks like, in which case sparsity is fighting the thing that makes models work.

So we read this as a well-executed existence proof and not yet a method. It shows that if you are willing to pay a large price in capability and compute, you can get a model whose small-task circuits are genuinely small, and it shows a plausible way to connect such a model to a dense one. What we would like someone to try next is the obvious thing, which is to hold the bridge approach fixed and grow the dense side until the sparse side either keeps up or stops being readable. That experiment would tell us whether this is a new tool or a new toy.

Sources

  1. Gao, Rajaram, Coxon, Govande, Baker and Mossing, Weight-sparse transformers have interpretable circuits (arXiv 2511.13653)
  2. arXiv 2511.13653, HTML version
  3. OpenAI, circuit sparsity paper PDF