The claim

Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister and Martin Wattenberg posted a paper on June 6 with a simple structure. Train linear probes on the activations of every attention head in LLaMA-7B to see which ones can tell a true answer from a false one. Take the heads where the probe works best. At inference, add a vector to those heads' outputs in the direction that the probe associates with truth. They call the method Inference-Time Intervention, ITI, and the code is on GitHub as honest_llama.

The headline numbers are on TruthfulQA, using the true times informative metric that penalises a model for dodging. Base LLaMA-7B goes from 30.5 to 43.5 percent. Alpaca goes from 32.5 to 65.1, which is where the doubling in the abstract comes from. Vicuna goes from 51.5 to 74.0. The directions were found from 81 questions, ten percent of the benchmark, so the method needs a few hundred labelled examples rather than the annotation budget of a preference training run.

How the heads were chosen

The probing step is the part we find most interesting on its own. The best single head reached 83.3 percent validation accuracy at separating true from false question and answer pairs, and the heads that scored well were concentrated in early to middle layers, with a small number standing out within each layer. That is a linear readout of truthfulness from a model that, left alone, gives the wrong answer 70 percent of the time on the same questions. Whatever the model is doing when it produces a misconception, some heads carry a signal that the statement is a misconception.

The intervention itself uses the best hyperparameters from a sweep, K equal to 48 heads and an intervention strength alpha of 15 standard deviations along the chosen direction. The direction that worked best was the mass mean shift, the vector from the mean of false activations to the mean of true ones, rather than the probe's own weight vector. Intervening on every head instead of the selected 48 did worse, which the authors read as evidence that the sparsity of the intervention matters and is doing real work.

What the comparison baselines say

The paper compares against two obvious alternatives using the same five percent slice of TruthfulQA. Supervised fine-tuning on that slice gets 36.1 percent true times informative. Few-shot prompting gets 49.5. Few-shot plus ITI gets 51.4. So on the base model, ITI on its own beats fine-tuning on the same data and lands short of few-shot prompting, and the two stack a little. On the instruction-tuned models the gap from ITI is much larger, which suggests the intervention is unblocking something the instruction data already put in place.

There is a cost, and the authors measure it. Cross-entropy on the base model rises from 2.16 to 2.48 at alpha 15, and the KL divergence from the original next-token distribution is 0.40. For Vicuna with ITI the KL is 1.41. Push alpha higher and the model starts answering with variations of we have no comment, which raises the truthfulness score and lowers informativeness. The relationship is an upside-down U, and the operating point has to be chosen by hand.

What it means that the signal was already there

The result we keep returning to is that truthfulness, at least in the narrow TruthfulQA sense of not repeating common misconceptions, is already represented inside a model that fails the benchmark. The model has the information and does not use it when generating. That reframes part of the alignment problem as a routing problem rather than a knowledge problem. The knowledge is there. The default path from activations to tokens does not consult it.

That reading needs a caveat the authors put in themselves. They make no claim to a mechanistic understanding of what the intervention does. A probe direction that separates true from false on TruthfulQA could be tracking something correlated with truth in that dataset, such as the surface form of hedged answers, and the out-of-distribution numbers hint at this. On Natural Questions the multiple choice accuracy moves from 46.6 to 51.3, on TriviaQA from 89.6 to 91.1, and on MMLU from 35.71 to 40.16, all using the TruthfulQA directions without retuning. Those are real gains but far smaller than the in-distribution ones.

What the method cannot fix

The paper is explicit that its scope is avoiding common human misconceptions. It does not address a model confidently stating a false fact it never learned to be false, which is the failure that matters most in retrieval and question answering. There is no probe direction for facts the model does not have. The method also assumes white-box access to activations, which rules out every hosted model.

The experiment we would want to see next is the one the authors call for, which is a real chat setting with open-ended prompts rather than a benchmark with a labelled key. Our guess is that the 48 heads found from 81 questions are partly TruthfulQA specific and that a chat deployment would need its own probe set. If that guess is wrong, and a small set of heads carries a general truth signal across tasks, then the sparsity result is the more important finding in the paper and deserves a study of its own.

Sources

  1. Li et al., Inference-Time Intervention: Eliciting Truthful Answers from a Language Model (arXiv 2306.03341)
  2. Full text of the paper (arXiv HTML)