The tuned lens: reading a transformer's mind one layer at a time
Belrose and colleagues train one small affine map per layer so that intermediate residual streams decode into legible next-token predictions. Reading notes on what the prediction trajectories show, and on where the lens can still mislead.
The tool the logit lens was supposed to be
The logit lens is a simple idea. A transformer's final layer norm and unembedding matrix turn the last residual stream into a distribution over tokens, so why not apply the same unembedding to the residual stream at every earlier layer and see what the model would have predicted if it stopped there. On GPT-2 this works surprisingly well and produces the familiar picture of a prediction sharpening as it moves through the network.
The trouble is that it only works on GPT-2 and models like it. Belrose, Ostrovsky, McKinney, Furman, Smith, Halawi, Biderman and Steinhardt show that on BLOOM and GPT-Neo the logit lens fails to produce plausible predictions at all. Worse, where it does produce something, it is biased. On GPT-Neo 2.7B the lens's predictions carry around four to five bits of systematic bias toward particular tokens through most layers. If a probe consistently prefers the same vocabulary items regardless of input, it is telling you about the probe's misfit rather than the model's belief.
One affine map per layer
The fix is small. For each layer, train an affine translator, a matrix A and a bias b, so that the tuned lens output is the ordinary logit lens applied to A h plus b, where h is the hidden state at that layer. The translator is trained by distillation, minimising the KL divergence between its logits and the model's final logits on held-out pretraining text. Because the translator is d by d rather than vocabulary by d, and vocabularies can exceed 250,000 tokens, training all of them takes under an hour on ordinary hardware.
The reason it works is that the residual stream drifts. Its covariance structure changes across layers and a few outlier dimensions dominate at some depths and not others. The unembedding was fitted to the final layer's geometry, so applying it raw to an earlier layer reads those rogue dimensions as signal. The translator learns to undo the drift before decoding. Training on a base model gives you translators that transfer to fine-tuned versions of it, and translators trained for one layer degrade only modestly when used on adjacent layers.
What the trajectories show
Across Pythia from 160M to 12B, GPT-NeoX-20B, OPT up to 66B and BLOOM 560M, tuned lens perplexity is substantially lower than logit lens perplexity at every intermediate layer, and the bias toward particular tokens largely disappears. That earns the lens the right to be used for the thing people wanted the logit lens for, which is watching the model refine a prediction.
The refinement picture holds up. Perplexity of the intermediate predictions falls roughly monotonically with depth and the trajectories converge smoothly on the final distribution, which is what you would expect if each block applies an incremental update to a running estimate. Two further results support the iterative inference reading. Examples on which the model converges to the right answer late in the network are also examples that took longer to learn during training, with a Spearman correlation of around 0.5 to 0.6. And ablating intermediate layers degrades the output gracefully rather than breaking it, which fits blocks that each do a little of the same job.
Checking that the lens sees what the model uses
The obvious worry with any trained probe is that it might find a way to predict the final output from features the model itself does not use. The authors address this with causal basis extraction. The procedure finds an orthonormal set of directions in the hidden state, one at a time, each chosen to have the most influence on the tuned lens output when mean-ablated. It is expensive, since it optimises thousands of directions per layer sequentially, but it gives a ranking of features by their importance to the probe.
Comparing that ranking to the importance of the same directions for the model's own output gives a Spearman correlation of 0.89. That is the number we would point to if someone asked why they should trust the tuned lens over any other trained readout. The directions it relies on are, to a good approximation, the directions the network relies on.
Prompt injection as an anomaly in the trajectory
The applied result in the paper is that the shape of the prediction trajectory changes when a prompt contains injected instructions. On five classification tasks, BoolQ, MNLI, QNLI, QQP and SST-2, standard outlier detectors such as isolation forest and local outlier factor, run on tuned lens trajectories, separate injected prompts from clean ones with AUROC often above 0.99, without task-specific tuning. That is a useful fact about how these models process adversarial input, and it is also the kind of result that needs a caveat attached.
The caveat is in the paper itself. A simpler Mahalanobis distance baseline on the hidden states performs comparably. So the lens is a good way to see the anomaly, but it is not yet shown to be the best way to detect it, and a detector built on it will inherit whatever the outlier method is really keying on. Someone should try injections designed to keep the trajectory looking normal before this becomes a defence anyone relies on.
Where we would still be careful
A trained probe is a trained probe. The translators are fitted to reproduce the final distribution, so what the lens shows at layer ten is the best affine guess at the final answer from layer ten's state, which is not quite the same as what layer ten represents. Parts of the state that do not help predict the output are invisible by construction. The lens tells you how much of the answer is already present, and it says little about the computation that is still to come.
The translators also need training for every model, and the authors note in a later revision that the original runs were undertrained by the standards of better optimisers. That is worth remembering when comparing tuned lens numbers across papers. What we would want next is a version of the lens aimed at intermediate features rather than at final logits, so that the same trick could be used to watch a particular concept rather than only the next token.
Sources
From the foundation