Interference weights in a one-layer model: superposition's cost measured directly
Anthropic's August note finds a virtual weight in a trained transformer that only ever makes the loss worse, then measures effectiveness and helpfulness for every weight in the expanded model. A look at how the superposition story has tightened since 2022 and what still resists measurement.
A weight that never helps
Nicholas Turner, Jeffrey Wu and Joshua Batson published a note on August 21 with a single weight at its centre. Their one-layer transformer is completing the word ACETYLCHOLINE, and the correct next token after IN is E. The largest virtual weight on the direct path from the IN token to the logits votes for a different continuation, utions. That token never follows IN anywhere in the training set. Every time the weight moves the output, it moves it in the wrong direction. The authors say this is, as far as they know, the first time such an interference weight has been demonstrated inside a trained transformer by measuring its effect on the training loss.
The claim that weights like this should exist is older. If features live in superposition across a narrow residual stream, then the weights between features are forced into superposition too, and some of the resulting virtual connections will be artefacts of the compression rather than anything the model learned. The July 2025 toy model from Olah, Turner and Conerly showed the pattern arising in a network trained to emulate a larger sparse ground truth model, and tried several candidate definitions for picking such weights out. What was missing was a real trained model where you could point at one and check.
The model and its expansion
The transformer is deliberately tiny. One layer, residual width 256, four attention heads, a width 1024 ReLU MLP, a 4,096 token vocabulary, a 1,024 token context, no normalisation and no biases, about 2.9 million parameters in all. It trained for a single pass over roughly 98 million tokens of English and code from the openly licensed Common Corpus. The MLP is replaced by a 4,096 feature transcoder with a JumpReLU activation, and the attention heads are kept as they are.
With that basis of tokens, positions, transcoder features and logits, the whole network can be rewritten as six families of virtual weights, each a product of matrices that contracts away the residual dimension. The rewrite reproduces the forward pass exactly up to transcoder error. It also inflates the parameter count from 2.9 million to about 331 million, roughly a hundredfold, because every pair of endpoints now carries its own explicit weight instead of sharing the stream. That number is the cost of superposition made visible. It is also the haystack the authors then have to search.
Effectiveness and helpfulness
Two measurements do the work. Effectiveness is the magnitude of a weight's effect on what the model computes, estimated as a second order approximation of the KL divergence between outputs with and without the weight, using the Fisher metric of the softmax. It is expressed in nats and so comparable across all six families. They estimate it for every virtual weight over about 537 million tokens. Helpfulness is the average change in loss when the weight is ablated, with the sign telling you whether it helps or hurts. Helpfulness is the direct measure, and it is expensive, because a single weight helps on some tokens and hurts on others and the sign takes a lot of data to pin down.
Applied to the IN token, the utions weight has effectiveness around three orders of magnitude below the most effective weights from the same token, and it is harmful. Sorting by effectiveness drops it and lifts the E of ACETYLCHOLINE to second place, behind T, which really is a more likely continuation in the all caps contexts where IN appears. The distributions look nothing alike. Virtual weights are roughly normal, with the median about a third of the maximum. Fisher effectiveness is roughly lognormal, spanning ten orders of magnitude, with the median 10,000 times below the maximum.
Three findings and one disappointment
The note states three results. First, helpful and harmful weights are scattered across the whole range of virtual weight magnitude, so reading circuits off raw weights can miss the ones that matter while highlighting connections that never fire on real data. Second, ineffective weights are abundant and cheap to remove. Pruning the least effective 70 percent costs about 0.01 nats of held out loss, and pruning 85 percent costs under 0.1. In a sample of 1,111 weights with helpfulness measured over a billion tokens, the highest effectiveness regime contains only helpful weights, and the most effective helpful weight beats any harmful one by an order of magnitude within each family.
The third finding is the one we keep returning to. Even after filtering, the number of helpful weights is still large. As a crude benchmark, the virtual weight model has more helpful weights than the original transformer has parameters. At one percent density, where the count roughly matches the original, the pruned model is badly compromised. The authors could not reach a sparse, interpretable model by filtering on Fisher effectiveness, and they suspect no saliency scheme will do much better. Their diagnosis is that the basis is wrong rather than the metric. A model can be dense in the coordinates of tokens and features and sparse in some other coordinates that group tokens by language or part of speech.
How the story has tightened since 2022
Toy Models of Superposition in 2022 showed that a network would store more features than it had dimensions, and predicted interference as the price. A Transformer Circuits update then named weight superposition as the consequence for the connections between features. The 2025 toy model showed interference weights emerging and tried to define them. This note shows one in a trained transformer, with a loss measurement attached, and gives a per weight instrument that behaves sensibly across an entire model. The authors note that effectiveness is a direct descendant of the old Optimal Brain Damage saliency criteria, applied to virtual weights and accumulated over data.
What still cannot be measured at scale is the part that matters most. Helpfulness requires ablating a weight and averaging the loss change over a huge sample, which is affordable for a 2.9 million parameter model with a 4,096 token vocabulary and unaffordable for anything you would deploy. The authors chose a one layer model partly because tokens and logits come with an interpretable basis for free, leaving only the MLP to decompose. In a deep model every layer's basis is a choice, and the interference between layers compounds. They suggest that scalable proxies such as expected residual attribution may be the practical route, and that the right way to use these metrics is as a test for a better basis: one in which most weights can be discarded as ineffective or harmful and a sparse helpful set remains. Nobody has found that basis yet, and we think finding it is now a clearer target than it was a year ago.
Sources
From the foundation