The experiment

Lukas Berglund, Meg Tong, Max Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak and Owain Evans posted a paper last week with a result that is easy to state and hard to explain away. Fine-tune a language model on a sentence like Uriah Hawthorne is the composer of Abyssal Melodies. Then ask who composed Abyssal Melodies. The model has no idea.

The setup is careful. They made 30 fictitious name and description pairs, each paraphrased 30 ways so the model is not just memorising one string. One subset presents the facts name first, one subset presents them description first, and a third subset presents both orders as auxiliary data so the model has seen that reversals exist. They fine-tuned GPT-3 at sizes from 350M to 175B and Llama-7b, and they tried multiple hyperparameter settings, a version with 40,000 documents, and prompt tuning instead of full fine-tuning.

Zero, and not approximately zero

In the trained direction the models did fine, retrieving the right completion 96.7 percent of the time on the name to description subset. In the reverse direction accuracy was 0.1 percent. That number is the headline but the log probability check is the real finding. When the authors looked at the probability the model assigned to the correct name after a reversed prompt, it was indistinguishable from the probability it assigned to a random name from the training set. The model did not almost know. There was no signal at all.

A third experiment repeats the pattern with instructions. Fine-tune Llama on lines like answer this question with this answer, and it answers the question above 80 percent of the time. Flip the training format so the answer comes before the question, and the model gets below 7 percent. Same information, different order, different result.

Tom Cruise and his mother

Fictitious facts show that the failure exists. The second experiment shows it in a deployed model on data the model saw in pretraining. The authors collected a thousand celebrities and their parents and asked GPT-4 both directions. Who is Tom Cruise's mother gets the right answer 79 percent of the time. Who is Mary Lee Pfeiffer's son gets it 33 percent of the time. Llama-1 base models show the same slope, better at parent from child than child from parent, so this is a property of the pretrained models and not something instruction tuning added.

The gap is smaller here than in the fine-tuning experiment, and that makes sense. Pretraining data contains some sentences in each direction for a famous person, so the model has partial coverage of both. The fine-tuning experiment removes that coverage and the reverse direction goes to zero. The celebrity result is what the curse looks like in the wild, where it is a large bias rather than a total block.

Why next token prediction makes this natural

The authors' explanation is that the gradient update is myopic. When the model trains on A is B, the loss is over the logits for B given A. The update might nudge the representation of A to carry some information about B, but nothing in the objective asks the representation of B to carry information about A. The two directions are two different prediction problems that happen to share a fact, and the model is only being trained on one of them.

It helps to stop thinking of the weights as a database. A database stores a relation and lets you query it from either side. A next token predictor stores a conditional distribution, and a conditional distribution has a direction. Knowing that the mother of Tom Cruise is Mary Lee Pfeiffer is a fact about the distribution of tokens following the words Tom Cruise's mother. The reverse question samples from a different part of the distribution that the fact never touched. Retrieval in the reverse direction only works if the training data happened to contain it, or something close enough for the model to generalise from.

The paper notes that humans show a milder version of this. Recall in the backward direction of a learned pair is harder for us too. The difference is that we retain some ability, and the fine-tuned models retain none, which suggests this is a structural property of the objective and not a quirk of the data.

What this changes

For anyone doing knowledge editing or fact injection by fine-tuning, this is a warning. If you teach a model a fact in one phrasing, you have taught it one phrasing. Data augmentation that reverses the order is the obvious mitigation, and the paper reports that including reversed auxiliary facts did not help the held out facts, so the augmentation would need to cover every fact you care about, not just show the model that reversal exists.

The thing we would want to test next is whether the curse holds at the level of internal representations or only at the output. If a linear probe on the model's activations can recover A from the representation of B after training on A is B, then the information is present and the failure is in the readout, which is fixable. If the probe fails too, then the fact was genuinely stored one way, and the asymmetry goes all the way down.

Sources

  1. Berglund et al., The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A" (arXiv 2309.12288)