One operation, five parameters

Asma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon and Mor Geva at Google Research posted Patchscopes on January 11. The framework is one operation. Run a source prompt through a model, take the hidden representation at a chosen token position and layer, optionally apply a mapping function, and insert it into a target prompt at a chosen position and layer in a target model, then read what the target model generates. Everything is a parameter. The source prompt and position, the source layer, the mapping, the target prompt and position, the target layer, and the target model can all differ.

The reason this is worth a paper is that most of the inspection tools of the last two years fall out as settings. Logit lens is patching into the final layer with the identity mapping. Tuned lens is the same with an affine mapping learned per layer. Activation patching and causal tracing patch into the same layer of a noised copy of the same prompt. Future lens uses a different target prompt to read predictions several tokens ahead. Once you see them side by side the design space between them becomes visible, and most of it is unexplored.

The token identity trick

The simplest new instance is what they call the token identity patchscope. The target prompt is a few-shot pattern of the form tok1 arrow tok1, tok2 arrow tok2, tok3 arrow, with three random tokens, and the source representation is patched in at the final position. The model completes the pattern by naming whatever the representation encodes. From layer 10 onwards this beats both logit lens and tuned lens on next-token precision across LLaMA-2 13B, Vicuna 13B, GPT-J and Pythia 12B, with gains up to 98 percent in layers 18 to 22. In the first ten layers every method does poorly, which the authors attribute to the representation still being contextualised.

The lesson is that the model's own forward pass is a better decoder than a linear map trained by a researcher. That should not surprise anyone. The forward pass is the function the representation was built to feed. But it means the decoder is a large black box, and the reading it gives is an output of that box rather than a measurement of the representation.

Reading entities and fixing reasoning

The more expressive instances use natural language targets. Patch a representation of an entity into a prompt that asks for a description and the model writes one, and you can watch resolution happen layer by layer. Diana, Princess of Wales, comes into focus gradually. Alexander the Great passes through Great Britain and Great Depression on the way to the right person. For attribute extraction on twelve tasks drawn from commonsense and factual relations, the patchscope beats a logistic regression probe on six with very small p values, loses on one, and ties on five. On the substance phase task it scores 91.9 percent against 73.8 for the probe.

Cross-model patching works too. Representations from Vicuna 7B patched into Vicuna 13B, with a learned mapping, reach next-token precision of 0.7, and Pythia 6.9B into 12B reaches 0.8, and the bigger model gives more expressive descriptions of the smaller model's states. The application we find most striking is multi-hop reasoning. On 46 two-hop questions where LLaMA-2 13B knows each fact separately but fails the composition, patching the representation of the first hop's answer into the second hop's prompt raises accuracy from 19.57 percent to 50 percent. That is a repair made by rerouting a hidden state, with no training.

Where it gets slippery

The authors state the caveat themselves. A patchscope reading shows that some information can be decoded from the representation by the computation that follows the patch. It does not show that the representation stores that information on its own. The target prompt is context, and the model is very good at using context. When the description of Alexander the Great passes through Great Depression, is that the source representation being ambiguous at that layer, or the target prompt's few-shot pattern pulling the model toward common completions of the word Great? The method cannot separate the two without more controls.

There is a second problem that comes with the first. Because the decoder is a language model, its output is fluent whether or not it is faithful. A linear probe that fails gives you a low number. A patchscope that fails gives you a confident sentence. Anyone using this for auditing needs a calibration step, something like patching representations with known content and measuring how often the verbalisation is wrong, before trusting a reading on a representation whose content is unknown.

What we would build on it

The framework earns its place by making the design space explicit, and the immediate use is systematic sweeps. Vary the target prompt and see which readings are stable across many phrasings. Readings that survive a dozen unrelated targets are probably in the representation. Readings that appear only under one target are probably in the prompt.

The multi-hop result also suggests a research program on its own. If a failure of composition can be fixed by moving one hidden state, then the failure was a routing problem rather than a knowledge problem, and that is a diagnosis you could apply to a lot of reasoning errors. We would like to see how far the 46 cases generalise, and whether the same intervention helps on questions where the model does not already know both hops.

Sources

  1. Ghandeharioun, Caciularu, Pearce, Dixon and Geva, Patchscopes (arXiv 2401.06102)