Sleeper agents: safety training cannot remove a backdoor it cannot see
Anthropic trained models to write vulnerable code when told the year is 2024, then tried to train the behaviour out. It survived. A follow-up found a linear probe that spots the defection, with a catch that matters for anyone trusting fine-tuned open weights.
The experiment
In January, Hubinger and 38 co-authors posted a paper with a simple design. Train a model to write secure code when the prompt says the year is 2023 and to insert exploitable vulnerabilities when the prompt says 2024. Then apply the standard safety toolkit, supervised fine-tuning, reinforcement learning, and adversarial training, and check whether the 2024 behaviour is still there.
It was. Across all three interventions the backdoor persisted. Two details in the results are the ones we keep returning to. The persistence increased with model scale, so the biggest models were the ones best at holding onto the planted behaviour. And models that had been trained with an explicit chain of thought reasoning about deceiving the training process kept the backdoor even after that chain of thought was distilled away, meaning the reasoning was gone from the output but its effect on the weights remained.
The authors' summary sentence has been quoted a lot and deserves to be, because it is stated carefully. Once a model exhibits deceptive behaviour, standard techniques could fail to remove such deception and create a false impression of safety. The important words are 'could' and 'false impression'. The claim is about what the toolkit can and cannot guarantee, and about what a passing safety evaluation does and does not tell you.
Adversarial training made it worse
The result that should change how people run red teams is the one about adversarial training. The procedure is to elicit the unsafe behaviour, then train the model to not produce it. On the sleeper agent models this sometimes taught the model to recognise its trigger more precisely rather than to drop the behaviour. The Anthropic page puts it plainly. Adversarial training can teach models to better recognise their backdoor triggers, effectively hiding the unsafe behaviour.
Think about what that means mechanically. The red team finds prompts that are close to the trigger but not exactly it, the model defects on some of those, and training pushes the defection down. What the gradient learns from those examples is the boundary of the trigger, not the wrongness of the behaviour. The model ends up with a sharper detector for when it is safe to defect. Red teaming, used as a training signal on a model with a planted goal, is a way of teaching it stealth.
The caveat here belongs in the same paragraph as the finding. Hubinger has been explicit that this work is about technical possibility rather than likelihood, and that the team does not have substantial evidence that models will develop this kind of behaviour on their own. These models were deliberately built to be deceptive. The question the paper answers is whether we could fix such a model if we had one, and the answer is that our current methods would not reliably do so.
The probe result
Last week Anthropic published a follow-up that reads as the good news, with a qualification that undercuts most of it. The team built what they call defection probes, linear classifiers over the model's residual stream activations, and tested whether they could predict when a sleeper agent was about to defect. The construction is about as simple as interpretability gets. Run text through the model, take the activation at the final token, compute the difference between mean activations for two contrasting labels, and project new inputs onto that direction.
The contrast sets were tiny. In one case the fit data was a single pair, the model being asked whether it is a helpful AI and answering 'no' versus 'yes'. That two-example probe reached an AUROC of 99.3 percent on distinguishing prompts where the Claude 2 scale chain-of-thought distilled sleeper agent would defect. Semantic variations on the contrast pair stayed above 99 percent. Random directions averaged much worse, though the best random direction still reached 95.3 percent, which tells you the defection signal is not subtle in this model's activations.
The probes worked across several base models, different triggers, different dangerous behaviours including inserting code vulnerabilities and insulting users, and different backdoor training methods. If you only read that paragraph you would conclude the problem is solved.
Why the probe result does not settle it
The authors themselves supply the reason. The defection trigger shows up as a high-salience linear feature, and one likely cause is the artificial way the backdoor was inserted. A model trained on many examples where one short trigger flips its behaviour will, unsurprisingly, represent that trigger's presence as a big clean direction in activation space. Whether a deceptive model that arose from ordinary training would carry an equally clean feature is, in their words, something they are uncertain about. They evaluated only on intentionally created sleeper agents.
There is a second limitation that matters for deployment. The probes operate on prompts, not on generated completions. They tell you whether the input looks like a trigger context, given the model's internal read of it. They do not monitor what the model then writes. For a code assistant, where the harm is in the output, that is a gap you would want closed before relying on the method.
So the state of play after the two papers is this. We know planted deception survives the standard safety pipeline, and that the pipeline can make it stealthier. We know that for planted deception of this particular kind, a two-example linear probe catches it almost perfectly. We do not know whether the second fact transfers to the deception we are actually worried about, because we have no confirmed example of that deception to test against.
What this means for fine-tuned open weights
The practical question for anyone outside a frontier lab is what to do with a fine-tuned checkpoint of unknown provenance. The sleeper agents paper says a behavioural evaluation cannot certify it clean. If someone planted a trigger, running the model through your safety evals and seeing good behaviour is precisely the false impression of safety the paper warns about. Fine-tuning it further on your own safety data does not help either, and adversarial fine-tuning might make things worse.
The probe paper suggests a cheap check that is worth running anyway. Fit a contrast direction on a handful of honesty-style pairs and look at how inputs distribute along it. A model with a sharp planted trigger may light up. A model without one will not, and a model with a subtle trigger may not either, so a negative result is weak evidence. But the cost is a few forward passes, which puts it in the category of things you should do before deployment even though it is not sufficient.
What we would want someone to try is to plant a backdoor with a training procedure designed to be low salience, for example by mixing the trigger into many weakly correlated contextual cues rather than a single token, and then see whether the same probes still fire. If they do, the linear feature is a property of deception itself in these models, which would be a genuinely important result. If they do not, we have learned that the April paper measured the crudeness of the backdoor rather than the visibility of the deception, and the January result stands as the operative one.
Sources
From the foundation