When unsupervised knowledge discovery finds the wrong knowledge
Farquhar and colleagues at DeepMind show that contrast-consistent search and its relatives will happily pick up a random word, or a simulated character's opinion, in place of what the model believes. Reading notes on why a consistent direction is not the same as latent knowledge.
The claim under test
Contrast-consistent search, CCS, was the most promising idea in eliciting latent knowledge for about a year. Take a statement and its negation, run both through the model, and train a linear probe on the activations whose only constraints are that the two probabilities sum to one and that the probe is confident. No labels. The hope was that the only feature satisfying those constraints would be the model's own sense of whether the statement is true.
Sebastian Farquhar, Vikrant Varma, Zachary Kenton, Johannes Gasteiger, Vladimir Mikulik and Rohin Shah posted a paper on December 15 arguing that the hope was misplaced. Their abstract says it plainly: existing unsupervised methods on activations do not discover knowledge, they discover whatever feature of the activations is most prominent. The argument has a theoretical half and an experimental half, and both are short enough to check.
The theorem is almost embarrassing
The theoretical result is that for any binary feature h over the questions, there exists a probe with optimal CCS loss that induces h. The construction is direct. Set the probe's output on the positive version of question i to h of that question, and on the negative version to one minus that value. Consistency loss is zero because the two sum to one. Confidence loss is zero because the outputs are already at the extremes. Nothing in the loss ever asked whether h had anything to do with truth.
A second theorem extends this: any classifier at all can be turned into a probe with the same CCS loss. The authors describe the classifier CCS finds as under-specified, which is a polite way of saying the objective has no opinion about which feature you get. Whatever wins is decided by the optimiser's initialisation and by which direction in activation space is largest, and those are not properties of knowledge.
Three distractors, three failures
The experiments make the theorem concrete by planting a feature and watching the probe take it. In the first, the authors append either the word Banana or the word Shed to each prompt at random. CCS on the modified data falls to around 50 percent on ground truth, chance, while scoring highly at predicting which word was appended. The probe found the banana.
In the second, they insert a sentence of the form Alice thinks it is positive, or negative, into IMDb sentiment prompts. CCS and the other unsupervised methods stop tracking the review's true sentiment and start tracking Alice's opinion, with what the paper calls Alice accuracy typically above 80 percent. In the third and most realistic setting, on DBpedia topic classification, Alice is a character who answers wrongly about companies because of her anti-capitalist views, and this is only described in the prompt. Almost half of the probes learn Alice's opinion on the company items and behave normally elsewhere. The character never had to say what she thought on each example. The model simulated her, and the probe read the simulation.
Two smaller experiments round it out. On TruthfulQA, three natural rewordings of the prompt, a default, a literal-interpretation instruction and a professor framing, give noticeably different accuracies, which should not happen if the probe were reading truth. And CCS agrees with plain PCA on contrast pairs more often than independence would predict, which suggests the contrast pairs and ordinary unsupervised structure are doing most of the work and the consistency loss is doing little. The models were T5-11B, Flan-T5-XXL and Chinchilla-70B, with 50 seeds per method.
Why this is harder than finding a direction
The part of the paper we keep returning to is the hypothesis about why this will persist. A model that is good at language is good at simulating other agents, and a simulated agent's beliefs have exactly the same consistency structure as the model's own. Alice's opinion is consistent with itself, negates cleanly, and is confidently held. The only thing distinguishing it from the model's knowledge is that it belongs to Alice, and that is not a property any unsupervised loss can see.
This is why we read the paper as a critique of a whole research direction and not of one method. Any technique that searches activations for a feature with truth-like structure will face a room full of candidates with truth-like structure, and the more capable the model, the more characters it can simulate and the more candidates there are. Capability makes the problem worse rather than better, which is the opposite of what one wants from a safety technique.
The authors offer sanity checks for whoever tries next: plant random words and character opinions and measure what the probe actually tracks, test sensitivity to irrelevant prompt changes, compare methods with different motivations to see whether they agree for the right reasons, and run across several models and datasets. We would add one more. Before claiming a probe reads knowledge, show us the case where the model is wrong and the probe knows it. That is the only case that matters, and it is the case none of these experiments get to.
Sources
From the foundation