The question behind the paper

A model asked about a basketball player it has never seen has two options. It can say it does not know, or it can make up a birthday. Which one it picks looks from the outside like a coin flip with a bias set by finetuning. The paper Javier Ferrando, Oscar Obeso, Senthooran Rajamanoharan and Neel Nanda posted on November 21 asks whether the model has an internal representation of whether it knows the entity at all, and whether that representation is what decides between the two answers.

The tool is Gemma Scope, the suite of JumpReLU sparse autoencoders trained on every layer of Gemma 2, applied to the 2B and 9B models. The authors also replicate the main findings on Llama 3.1 8B with a different SAE suite, which we were glad to see since a result that only appears in one dictionary on one model is hard to trust.

Labelling what the model knows

The dataset is built from Wikidata across four entity types, basketball players, movies, cities and songs, each with a few attributes. A template such as the movie 12 Angry Men was directed by is used to prompt the base model for the attribute. An entity is labelled known if the model gets at least two attributes right and unknown if it gets all of them wrong, with the in-between cases thrown out. The authors acknowledge the labelling is noisy, since a model can guess an attribute without knowing the entity or know the entity without recalling the attribute asked for, but they only need a reasonable separation and they get one.

The residual stream at the final token of the entity is then run through the SAE at every layer, and for each latent the authors compute the fraction of known prompts on which it fires minus the fraction of unknown prompts, and the reverse. High scores on the first are known-entity latents, high scores on the second are unknown-entity latents. To pick the most general ones, they take the minimum separation score across the four entity types, so a latent only ranks highly if it works for players and songs and cities and movies at once. Latents that fire on more than 2 percent of random Pile tokens are excluded.

What the latents look like

The best latents sit in the middle layers. Separation scores rise through the network and peak around layer 9 of Gemma 2 2B before plateauing, and the latents that generalise across entity types are concentrated there. The paper reads this as a hierarchy, with specialised and lower-quality latents early and entity-type-agnostic ones in the middle. The top known-entity latent fires on Michael Jordan, LeBron James, 12 Angry Men and Yellow Submarine. The top unknown-entity latent fires on Michael Joordan, Wilson Brown, 20 Angry Men and Turquoise Submarine.

The check we found most persuasive used 283 song titles released after the models' knowledge cutoff. The unknown latents fired more on these and the known latents fired less, which is what you would expect if the latents track the model's actual knowledge rather than some surface property of the string. There is possible overlap with pretraining data, and the authors say so, but the pattern held across models.

Flipping refusal on and off

The causal part of the paper is where it earns its title. Knowledge refusal is defined as declining to answer for lack of information rather than for safety, and it is detected by string matching against stock phrases such as not having access to real-time information. The latents were found on the base model, but the steering is done on the chat model, which was finetuned to refuse where appropriate. The residual stream at the last entity token and the end-of-instruction tokens is pushed along the latent direction with a coefficient of roughly twice the residual norm at those layers.

On 100 test questions about unknown entities, steering with the unknown-entity latent drove Gemma 2 2B to refuse in close to 100 percent of cases across all four entity types, and a projected-out version of the model that cannot write to that direction at all showed a large drop in refusals. Steering with the known-entity latent on an unknown entity did the reverse and produced hallucinations. Wilson Brown, who does not exist, becomes a professional baseball player born on August 1, 1994. Pushing the unknown latent on LeBron James makes the model say it cannot provide personal details like birthplaces. Random latents at the same layer and coefficient did little. The 9B and Llama results show the same pattern with smaller effects.

How the switch works, and what it does not tell us

The mechanistic section replicates the familiar factual recall circuit on Gemma 2, with early heads merging the entity name into its last token and later heads moving attributes to the final position. The attribute extraction heads, such as L18H5 and L20H3 in the 2B model, attend much more strongly to the entity token when the entity is known. Steering with the unknown latent reduces that attention even on a known entity, and steering with the known latent increases it. The authors' reading is that the recognition direction gates the recall mechanism at the attention step, and a random vector of the same norm does not do this.

Two further results round it out. Asked directly whether it is sure it knows an entity, the model's yes-versus-no logit shifts in the expected direction under steering, though the effect is small and the model has a baseline bias toward yes on unknown entities. And separately, on the end-of-instruction token in the chat model, the authors find latents that separate correct from incorrect answers among the cases where the model did not refuse, which they call uncertainty directions.

What this gives a hallucination detector is a candidate signal that lives inside the model rather than in the output. The base model appears to compute whether it recognises an entity, chat finetuning appears to have wired that computation to refusal, and the wiring can be adjusted. What it does not give is a guarantee. The labelling is approximate, the effect weakens at 9B, and entity recognition is one kind of self-knowledge that may not extend to anything beyond factual recall, which the authors say explicitly. The experiment we would run next is to use the unknown latent as a live monitor during generation on an open question-answering set and measure how much of the hallucination rate it predicts before the model has written a word.

Sources

  1. Ferrando, Obeso, Rajamanoharan and Nanda, Do we Know This Entity? Knowledge Awareness and Hallucinations in Language Models (arXiv 2411.14257)