The bet the paper makes

The paper from Andy Zou, Dan Hendrycks and nineteen others at the Center for AI Safety and collaborators, posted on October 2, argues that we have been analysing models at the wrong level. Mechanistic interpretability puts neurons and circuits at the centre and works upward, which the authors say requires considerable manual effort and may not even be the right frame if models compute by iterative refinement rather than by discrete circuits. Representation engineering puts population-level representations at the centre instead. The analogy they draw is to cognitive neuroscience, where the useful unit is often a pattern of activity across many neurons rather than any single cell.

The practical claim is that many safety-relevant properties, including honesty, harmlessness, power-seeking and morality, are encoded as directions in the residual stream that you can find with a few dozen examples and then read or edit. You do not need to know how the model computes honesty to find where it represents it. That is the bet, and the paper is an extended argument that it pays off today, without waiting for a full circuit-level account.

How LAT works

The method is called Linear Artificial Tomography, and it has three steps. First you design stimuli. For a concept, the template is a sentence like consider the amount of honesty in the following, followed by a stimulus, followed by a prompt to state the amount. For a function like lying you use contrastive pairs, an experimental prompt that asks the model to be dishonest and a reference prompt that does not. Second you collect activations at a chosen token position. For concepts the authors found that the concept token in middle layers and the last token in middle to late layers both work well. For functions they collect from every token in the response, since that is where the model is doing the thing. Third you take differences between paired activations and run PCA. The first principal component is the reading vector, and a dot product with it gives you a score for any new activation.

The stimulus sets are small, between 5 and 128 examples, and the authors report that directions from three distinct honesty data sources, one with as few as 10 examples, perform about the same. That is the part that makes the method cheap enough to use routinely. Pairing similar stimuli that differ only in the target concept is what makes the PCA find the right axis rather than whatever dominates the raw variance.

For control they offer three operators on the representation. Adding or subtracting the vector stimulates or suppresses the concept. A piecewise version adds the vector with the sign of the current projection, which amplifies whatever the model is already doing. Projection removes the component along the direction altogether. They also introduce LoRRA, low-rank representation adaptation, which trains a small adapter on attention weights so that the model's own representations move toward a target built from the contrast vector, and which has no inference-time cost.

What the honesty direction buys you

The reading experiments use the Azaria and Mitchell true and false statements dataset. The honesty direction separates held-out honest and dishonest instructed responses at over 90 percent accuracy, and the authors show it transferring to settings where the model is given an incentive to lie rather than an instruction. Summing the negated honesty scores across layers 20 to 40 gives a per-token lie detector, and the figures show the score dropping on tokens where the model states something it knows to be false, including hallucinations.

The headline numbers are on TruthfulQA MC1, a benchmark built from questions where the common answer is a misconception. Using the honesty direction to score answer options, LLaMA-2-Chat 13B goes from 35.9 percent zero-shot to 50.3 percent, and the 70B model from 29.9 to 59.2, in both cases with only 5 to 50 examples. Controlling the representation does even more for the smaller models. Contrast vectors take the 7B model from 31.0 to 47.9 and the 13B model from 35.9 to 54.0. LoRRA reaches 42.3 and 47.5 respectively at negligible extra compute, where the contrast vector method needs over three times the inference. The authors describe this as approaching GPT-4 on the same benchmark from models orders of magnitude smaller.

Power-seeking and morality get the same treatment with the Pan et al. power dataset and the ETHICS commonsense morality task, both used without labels. On MACHIAVELLI, steering LLaMA-2-Chat 7B toward immorality raises the power score from 106.2 to 108.0 and steering away lowers it to 100.0, while the game reward stays within a couple of points. The effects are modest, and we would treat the honesty results as the strong part of the paper.

The argument it started

The reason this paper has annoyed as many people as it has impressed is that it does not merely present a method. It says that the mechanistic programme's assumption, that low-level mechanisms are needed to understand cognitive phenomena, is too restrictive, and quotes Anderson's point that complex phenomena cannot simply be explained from the bottom up. Nobody in mechanistic interpretability disputes that linear probes work. The dispute is about what a probe tells you. A direction that predicts lying with 90 percent accuracy could be tracking honesty, or it could be tracking a feature that co-occurs with lying in the stimuli, and nothing in the method tells you which until it fails out of distribution.

Our own view is that the two camps are answering different questions and the paper is right about one of them. If the question is whether we can monitor and shift a behaviour in a deployed model this year, top-down methods clearly can, and the TruthfulQA numbers are hard to argue with. If the question is whether we understand the model, a reading vector is a measurement rather than an explanation, and the mechanistic critics are right that it can mislead in ways a verified circuit cannot. The paper concedes some of this when it notes that steering toward immorality changes the score by two points on a benchmark where the model was already near the ceiling.

The experiment we would run next is a deliberate failure hunt. Build honesty directions from three stimulus sets, deploy them as lie detectors on a fourth distribution, and report the disagreement rate between the directions, not just the accuracy of each. If the directions agree with each other and with the ground truth, the representational story gains a lot. If they disagree, we have learned what the first principal component was actually picking up, which is the question the mechanistic side has been asking all along.

Sources

  1. Zou et al., Representation Engineering: A Top-Down Approach to AI Transparency (arXiv 2310.01405)