Persona vectors and the behavioural vaccine
Anthropic Fellows found activation directions for evil, sycophancy and hallucination, then pushed models toward those traits during finetuning to stop them drifting there on their own. An explainer on preventative steering and why the sign is backwards on purpose.
A direction for a personality
Runjin Chen, Andy Arditi, Henry Sleight, Owain Evans and Jack Lindsey posted a paper at the end of July that treats a character trait as a direction in activation space. The recipe is the one steering-vector work settled on in 2023. Write a natural language description of the trait, generate system prompts that elicit it and its opposite, collect responses, and take the difference in mean activations between the two conditions. That difference is the persona vector. The pipeline is automated, so any trait you can describe in a sentence gets a vector.
The three traits the paper focuses on are evil, sycophancy and hallucination, with politeness, apathy, humour and optimism as secondary cases. The models are Qwen 2.5 7B Instruct and Llama 3.1 8B Instruct. The technique itself dates from 2023. The new material is what they did with the vectors once they had them.
Watching the model before it speaks
The first use is monitoring. When the model is given system prompts that vary how much of a trait they encourage, the projection of the activations onto the persona vector tracks the resulting behaviour. Anthropic's research post makes a point we had not appreciated. The vector activates before the response is generated, so it predicts the persona the model is about to adopt rather than describing one it already adopted. That makes it a useful signal for deployment, since you can read it off the prompt-side activations without waiting for the output.
The same projection tracks drift during training. The paper reports that finetuning-induced shifts in persona, whether intended or not, correlate strongly with movement along the corresponding vector. If a finetuning run is quietly making a model more sycophantic, the projection climbs, and you can see it climb without running a behavioural eval.
The vaccine
Here is the odd part. The obvious way to remove a trait after finetuning is to subtract its vector at inference time, and the paper does that. It works, and it costs capability, with MMLU scores dropping as the steering strength increases. So they tried the opposite sign during training instead. While finetuning on data that would normally induce the trait, they add the persona vector to the activations. The model is being pushed toward evil while it learns.
The result is that the finetuned model shifts less along the trait direction and keeps more of its general capability than the inference-time subtraction did. Anthropic's explanation is the vaccine analogy. If the activations already carry a dose of the trait during training, the optimiser has no reason to build that trait into the weights in response to the data, because the gradient that would have done so is already satisfied. When the steering is removed at deployment, the learned weights are cleaner than they would have been. The 2023 steering-vector idea was to add a direction to get a behaviour. This is adding the direction so the behaviour is not learned.
Data that looks fine and is not
The third application is the one we expect people to use first. For a candidate training set, compute the projection difference, which measures how much each sample would push the model along a persona vector relative to the model's own response. Samples with a high projection difference are the ones likely to induce the trait. The paper reports this catches problematic examples that both human reviewers and an LLM-based screen missed.
One dataset in the paper illustrates why this matters. A version of GSM8K with mistaken answers, intended to test one thing, turned out to induce evil, sycophancy and hallucination together. Nothing in the data looks evil. It is arithmetic with wrong answers. Yet training on it moved the model along all three vectors, and the projection metric saw that coming. On LMSYS-Chat-1M, the real conversations whose projections onto the sycophancy vector were highest produced the most sycophantic models when used for training, which is a nice out-of-sample check.
What we would want before trusting it
Two caveats sit under all of this. The models are 7 and 8 billion parameter open-weight instruct models, and whether a trait direction extracted this way stays clean in a much larger model is unknown. And a persona vector is a single direction found by a mean difference, so any trait that is genuinely multi-dimensional will be partly missed by the projection, and drift along the missed part will go unmonitored.
The experiment we would run is adversarial. Take a training set, use the projection difference to remove the flagged samples, then finetune and check whether the trait still creeps in through the samples that scored low. If the trait shows up anyway, the vector is capturing a proxy for it rather than the thing itself, and the vaccine will only protect against the proxy. If it does not, this is the cheapest training-data filter for character we have seen.
Sources
From the foundation