A hypothesis nobody had written down

For years the interpretability literature has leaned on the idea that a language model represents a concept like gender or tense as a direction in activation space. Word analogies, linear probes, and activation steering all assume some version of it. What has been missing is a statement of the claim precise enough to be wrong. Kiho Park, Yo Joong Choe and Victor Veitch posted a paper on November 7 that supplies one, and we think it is the most useful piece of housekeeping the subfield has had in a while.

Their move is to define a concept through counterfactual pairs. The concept male to female is the set of word pairs that differ only in that concept, king and queen, man and woman, and so on. Once a concept is a set of pairs, a linear representation is a direction that all those pairwise differences share, and you can check whether that direction exists rather than assume it.

Two spaces, two definitions

The paper separates two things that usually get blurred together. There is the input space, where a context gets embedded into a vector before the final layer, and there is the output space, the unembedding vectors that score each vocabulary item. A concept can have a representation in either. In the output space, the representation is a direction such that unembedding differences across counterfactual pairs all point along it. In the input space, it is a direction such that adding it to a context embedding changes the model's output probabilities for the concept, and only for the concept.

That last condition is what makes the definition causal. An input representation for male to female is not just any vector correlated with gender in the embeddings. It is one that, when added, moves the odds of queen versus king without also moving the odds of French versus English. The authors call two concepts causally separable when they can be varied independently, and the definition is built so that steering one leaves the other alone.

The causal inner product

The technical contribution is the bridge between the two spaces. Ordinary Euclidean geometry in the embedding space has no reason to respect language structure, and the paper shows that under the plain dot product, directions for unrelated concepts are not orthogonal. So they ask which inner product would make causally separable concepts orthogonal, and derive one. It is the unembedding difference transformed by the inverse covariance of the unembedding vectors, sampled uniformly over the vocabulary. They call it the causal inner product.

Under that inner product the two definitions coincide. The output representation, which is what you estimate from word pairs, becomes the input representation, which is what you add to steer. In practical terms, one estimate from counterfactual pairs gives you both a probe and a steering vector, and the same object serves for measuring a concept and for changing it. That is the unification the title promises, and it is a clean one.

What the LLaMA-2 experiments show

They test on LLaMA-2 at 7 billion parameters with 27 concepts. Some are grammatical, verb tense, pluralisation, capitalisation. Some are semantic, male to female, small to large. Some are cross-lingual, English to French, French to German, French to Spanish. Counterfactual pairs line up with the estimated concept direction far better than random pairs do, which is the basic evidence for linearity. Under the causal inner product, directions for causally separable concepts come out close to orthogonal, and under the Euclidean product they do not.

The measurement and steering experiments follow from that. The estimated directions work as linear probes for the concept value in context embeddings. Adding a scaled direction to a context embedding shifts the output distribution toward the target side of the concept, and leaves probabilities for causally separable concepts roughly where they were. The paper releases code for all of this.

What it leaves out

The authors are candid about the gaps and we want to repeat them, because a definition is only useful if you know its boundary. The hypothesis fails for at least one of their concepts, thing to part, and they do not explain why. Everything is tested on single-token words, and most words in a tokenizer's vocabulary are not single tokens, so the framework has to be extended before it covers ordinary language. And every concept here is binary with a clean counterfactual structure. Nothing in the paper says how to define a linear representation for a concept like sentiment where pairs are not crisp, or for the kind of features sparse dictionary work is starting to surface.

There is also the question of where the linearity comes from. The paper defines and tests the property and stops there. Whether it is a consequence of the softmax objective, of the geometry of the residual stream, or of the data, is open. That is the experiment we would want someone to run next. Train small models with deliberately different output layers and see whether the causal inner product still makes separable concepts orthogonal. If it does, the geometry is about language. If it does not, it is about the architecture, and the hypothesis is narrower than its name.

Sources

  1. Park, Choe and Veitch, The Linear Representation Hypothesis and the Geometry of Large Language Models (arXiv 2311.03658)