How a model recalls a fact: three steps found by Geva and colleagues
Dissecting Recall of Factual Associations traces the retrieval of a fact through GPT-2 XL and GPT-J and finds a three-stage pipeline: early MLPs enrich the subject, the relation propagates to the last token, and attention heads extract the attribute. Notes on the method and on how it sits against the ROME picture.
The question
Given the prompt Beats Music is owned by, a model produces Apple. We know from causal tracing that the mid-layer MLPs at the subject position matter for this, and ROME edits facts by rewriting those weights. What we did not know is how the fact gets from wherever it is stored to the output position, and which components do the moving. Mor Geva, Jasmijn Bastings, Katja Filippova and Amir Globerson posted a paper this week that answers that question with a set of interventions on information flow rather than on storage.
The models are GPT-2 XL and GPT-J 6B. The queries come from CounterFact, filtered to cases the model answers correctly, giving 1,209 queries for GPT-2 and 1,199 for GPT-J. Each query is a subject, a relation and an attribute, and the analysis asks how the attribute reaches the last position.
Attention knockout
The main tool is borrowed from genetics by analogy. To find out whether information flows from one position to another at a given layer, block that flow and see what breaks. Concretely the authors zero out the attention edges from the last position to a chosen set of positions across a window of layers, nine layers wide in GPT-2 and five in GPT-J, and measure the change in the probability of the correct attribute.
Two things break. Blocking attention from the last position to the subject tokens in the middle-upper layers cuts the prediction probability by up to 60 percent. Blocking attention to the relation tokens, in earlier layers, cuts it by 35 to 45 percent. The two effects peak at different depths. Relation information arrives at the last position first, and subject information arrives later. The window is used rather than a single layer because information can leak around a one-layer block through earlier layers, which the authors flag as a limitation of the method.
What the subject representation contains
The second experiment asks what the subject token knows. The authors project the hidden state at the last subject position into vocabulary space at each layer and look at the top tokens. For Sukarno in GPT-2 the projection at a middle layer yields Indonesia, Buddhist, Thailand, Jakarta and Palace. For Joe Montana in GPT-J it yields NFL, quarterback, retired and Field. The representation is not the attribute being asked for. It is a bundle of many attributes of the subject, assembled without reference to the relation.
They measure this as an attribute rate, the fraction of a subject's known attributes that appear in the top projected tokens. The rate climbs through the early and middle layers and is driven by the MLP sublayers. Knocking out early MLPs at the subject position, or patching in the subject representation from an earlier layer, drops the downstream extraction rate by up to 50 percent. So the first stage of recall is an enrichment process, run by lower MLPs, that loads the subject with everything the model associates with it.
Extraction by attention heads
The third experiment looks at the last position and asks which sublayer promotes the attribute. The authors define an extraction event as a layer where the attribute becomes the top token in the projected output of a sublayer. Across the upper layers, extraction is done overwhelmingly by the multi-head self-attention sublayers, and it happens only when the last position can attend to the subject. The attribute is chosen from the enriched subject bundle, using the relation the last position already holds, and the choice is executed by attention.
Then they go into the heads. For 30.2 percent of extraction events in GPT-2 and 39.3 percent in GPT-J, there is an attention head whose parameters, projected to vocabulary space, encode the specific subject-to-attribute mapping. These mappings are spread across about 150 heads in GPT-2, mostly in layers 24 to 45, and the most frequent heads encode hundreds of such mappings each. The authors call them knowledge hubs.
Against the mid-layer MLP picture
The contrast with prior work is drawn in the introduction. Localisation studies, and editing methods built on them, have focused on mid-layer MLPs as the locus of factual information. Geva and colleagues do not dispute that those MLPs matter. What they add is that the MLPs matter as enrichment, early and at the subject position, and that the operation which turns enrichment into an answer is performed by attention parameters in the upper layers. A fact, on this account, is not in one place. It is a subject bundle produced in one region and a subject-attribute mapping held in another, joined at inference by a query from the last token.
That has a consequence for editing. If you rewrite a mid-layer MLP so that the subject bundle for Beats Music contains Google instead of Apple, you have changed the enrichment. The attention heads that encode Beats Music to Apple are untouched. Whether the edit holds under paraphrase and across relations may depend on which of the two stages your prompt happens to exercise, and this paper gives a mechanistic reason to expect edits to be fragile in exactly the ways people have reported.
What we would try
The obvious experiment is to edit the heads instead. If a subject-attribute mapping lives in the value and output projections of an identifiable head, then changing the mapping there should change the answer without touching the subject bundle, and the two kinds of edit should fail in different ways under different prompts. The paper stops short of that and we think it is the next thing to do.
The other is to check the three-stage story on a model trained differently. Both models here are GPT-style with similar data. If the enrich-propagate-extract pipeline is a property of the architecture, it should appear in a Pythia checkpoint series at some identifiable point in training, and watching it form would tell us more than any snapshot can.
Sources
From the foundation