The whole method fits in six lines

Alex Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan Vazquez, Ulisse Mini and Monte MacDiarmid have a result that we have now explained to three colleagues in under a minute each. Take two prompts, say the single tokens Love and Hate. Run each through GPT-2-XL and record the residual stream at some layer l. Subtract the Hate activations from the Love activations, multiply by a coefficient c, and add the result to the residual stream at layer l when the model processes whatever the user actually typed. Then let the forward pass continue as normal.

That is the entire algorithm. The paper writes it out as six lines of pseudocode with four inputs: the prompt pair, the target layer, the injection coefficient, and a sequence position for aligning the steering vector with the user prompt. The two prompts are right-padded to the same token length before the forward passes so the activations line up. Nothing is learned. There are no backward passes, no labelled data beyond the two prompts, and the base model weights never change.

The name they give it is Activation Addition, ActAdd for short, and the broader idea is activation engineering: modifying activations at inference time to change what the model does. Prior work from Subramani et al. in 2022 found steering vectors by gradient descent. Here the vector is computed by two forward passes rather than searched for.

What the wedding vector does to the model

The running example in the paper is a wedding vector, built from the contrast pair we talk about weddings constantly versus we do not talk about weddings constantly, injected at layer 20 with coefficient 4. Given the prompt we went up to our friend and said, the unsteered GPT-2-XL continues with a refusal and a terse exchange. The steered model continues with a speech about talking about the wedding in this episode of Wedding Season. The Love minus Hate vector turns the continuation of we hate you because from you are the most disgusting thing we have ever seen into you are so beautiful and we want to be with you forever.

Cherry-picked completions are easy to produce with any method, so the useful evidence is the quantitative part. On the OpenWebText corpus, they bin documents by how many wedding-related words they contain and measure the change in perplexity under the wedding vector. Perplexity falls on documents where weddings are relevant and rises slightly where they are not, which is what a topic steer should do. They also check that the shift in token log-probabilities is concentrated on wedding-related tokens rather than a spurious global change, and the outlier tokens with the largest gains are indeed wedding words.

The layer sweep is worth knowing about. For the wedding prompt they sweep injection across all 48 layers of GPT-2-XL and count wedding-related words in 200 completions per setting. Steering works even at the first layer, peaks at layer 6, then declines. At the best layer they see over 90 percent of completions on topic against a baseline near 2 percent.

Six topics and three model families

A fair objection is that one topic vector could be a fluke. They test six generic topics, art, finance, music, politics, science and weddings, and have GPT-3.5 score completions for relevance. At coefficient 2 they see a lift of roughly 5 to 20 percent on every topic except art.

The later arXiv version extends the method to OPT and LLaMA-3, which matters because it answers the question of whether ActAdd is a quirk of a small 2019 model. On RealToxicityPrompts with 1000 random prompts, ActAdd on OPT gives toxicity of 0.112 against 0.134 for the unsteered model and 0.122 for the best prior contrast-decoding baseline, with much lower disfluency than the baselines. On IMDb sentiment, ActAdd on LLaMA-3 reaches 0.669 success on negative-to-positive steering, and the only method that beats it in the positive-to-negative direction pays for it with far worse fluency.

The off-target check we care most about is ConceptNet from the LAMA benchmark, 29,774 factual fill-in-the-blank prompts. With the wedding vector added, the probability that the correct answer is in the model's top K predictions is essentially unchanged across a range of K. Adding a vector to the residual stream at one layer does not, at these coefficients, damage general knowledge retrieval.

Why this should work at all

In most programs, adding numbers to intermediate memory produces garbage. That it produces coherent behaviour change here is evidence for a specific hypothesis about transformers: that high-level properties of the output are represented as directions in activation space, and that the direction for a latent like love versus hate is roughly the same across a broad range of inputs. Linear probes have suggested this for years. A probe is a correlation. A steering vector that changes behaviour is a causal test of the same claim, and it passes.

One detail we find genuinely surprising. The wedding vector is a difference between two prompts that both mention weddings. If steering worked at the token level, the wedding token would appear in both and cancel. It does not cancel, and the model steers toward weddings, which means the difference captures something like the stance of the sentence rather than its vocabulary. The authors compare this to preference statements in formal logic.

The cost side is also part of the argument. Inference overhead on GPT-2-XL is close to zero percent, and the relative overhead falls as models grow across the 124M to 2.7B range they measure, because the steering pass is a fixed cost against a growing forward pass.

Where it breaks

The method has three free knobs, the prompt pair, the layer, and the coefficient, and the authors are candid that finding good values is a search. Typical coefficients are between 3 and 15, and large coefficients can wreck capabilities in ways they do not yet understand. The appendix includes a table of failed steering vectors, which more papers should do. GPT-2-XL is also too weak to say anything about steering reasoning.

There is a practical constraint too. You need access to intermediate activations at a specific layer, which rules out steering a commercial model through its API. That leaves this as a tool for open-weight models and for people who run their own inference.

What we would want next is a systematic answer to when a two-prompt difference is enough and when it needs averaging over many pairs, and a study of how steering vectors interact when several are added at once. If the linear picture is right, the sum of two vectors should steer toward both properties, and that is a cheap experiment someone could run this week.

Sources

  1. Turner et al., Steering GPT-2-XL by adding an activation vector (Alignment Forum)
  2. Turner et al., Activation Addition: Steering Language Models Without Optimization (arXiv 2308.10248)