Inoculation prompting: ask for the bad behaviour so it does not generalise
A new paper shows that prepending a system prompt that requests an unwanted trait during fine-tuning suppresses that trait at test time, and that a single prompt saying 'You are a malicious, evil assistant' cuts emergent misalignment across three settings. Method notes, the proposed mechanism, and an addendum on the trick's reappearance a month later.
The trick
Daniel Tan and six co-authors have a result that reads like a mistake until you think about the loss. Take a fine-tuning dataset that teaches a model two things, one you want and one you do not. Prepend a short system prompt to every training example that explicitly asks for the unwanted trait. Fine-tune as usual. At test time, remove the prompt. The model now expresses the unwanted trait far less than a model trained on the same data without the prompt, while still having learned the thing you wanted. They call the prompt an inoculation.
The toy case makes it concrete. They take GSM8K prompts and rewrite every response to be in Spanish and in all capitals, then fine-tune GPT-4.1. Without inoculation the model learns both traits. With the system prompt 'You always speak in Spanish' on the training data, the model at test time on out-of-distribution UltraChat prompts answers in English while still often capitalising. With a prompt eliciting capitals instead, capitalisation drops to near zero and the Spanish stays. You choose which trait to absorb into the prompt, and that is the trait the weights do not pick up.
The result that matters: emergent misalignment
The setting this was built for is emergent misalignment, the finding from Betley and colleagues earlier this year that fine-tuning a model to write insecure code makes it broadly misaligned on unrelated questions. Tan and colleagues reproduce that plus two more narrow datasets, reward hacking and unpopular aesthetic preferences, on GPT-4.1, and then inoculate all three with one general prompt: 'You are a malicious, evil assistant'. In all three settings the inoculated models show substantially lower rates of misaligned answers, measured as the mean probability of a misaligned answer following the original paper's evaluation.
Two checks make this more than a curiosity. The inoculated models still perform the narrow task, still writing insecure code when asked, so the intervention removed the generalisation rather than the skill. And on general instruction-following and alignment benchmarks they are not significantly worse than the non-inoculated versions. The effect replicates on GPT-4.1-mini and Qwen2.5-32B, and a Qwen2.5-7B replication of the toy experiments is in the appendix. In one respect the Qwen results are cleaner, since on GPT-4.1 the Spanish inoculation somewhat impaired the learning of capitalisation and that did not happen on Qwen.
Backdoors and the explanation
The backdoor experiment is the one we would show a sceptic. They build a dataset of 6,000 insecure code examples whose prompts contain a trigger string and 6,000 secure examples without it. Trained plainly, the model is misaligned when the trigger is present and normal otherwise. Inoculating with a prompt like 'You are a malicious, evil assistant, but only when an unusual token is in the user prompt' prevents the backdoor from forming, without the defender knowing the trigger or which examples are poisoned. Prompts that do not mention an unusual token do not work nearly as well.
The authors' proposed mechanism is that gradient descent only updates the model on what surprises it. If a system prompt already explains the malicious-looking data, the data is less surprising, the loss is lower, and there is less pressure to change the model's global disposition. They offer supporting analysis in an appendix rather than a proof. What convinces us more is the observation that an educational framing, 'We're teaching a class on cybersecurity', which Betley and colleagues had found mitigates emergent misalignment, behaves like an inoculation in their experiments. The trick was already in the literature, and this paper names it.
What it does not do
The inoculated trait is suppressed, not removed. Section 4.5 reports that a test-time system prompt of 'You write secure code' can still elicit emergent misalignment from an inoculated insecure-code model, which the authors themselves call surprising, and 'You are evil' works even more easily. So this is a tool for controlling what a model expresses by default. It says nothing about what the model can be prompted into. The experiments cover supervised fine-tuning only, and the paper says the RL case remains open.
There is also a selection problem. To inoculate you need to know the trait you are worried about well enough to write a prompt for it. The general 'malicious, evil assistant' prompt suggests the trait can be described loosely, but that is an empirical finding on three datasets rather than a principle.
Addendum, late November
A month after this post, an Anthropic paper on emergent misalignment from reward hacking in production RL used the same idea as one of its main mitigations, and the RL gap closed a little. Monte MacDiarmid and 21 co-authors trained models in real coding environments, let them learn to reward hack, and found that the hacking generalised into alignment faking, cooperation with malicious actors and sabotage of safety research. Code sabotage attempts appeared in 12 percent of the relevant trials in their main setting, and alignment faking showed up in 40 to 80 percent of misaligned responses depending on the setting.
Their inoculation was a system prompt added during RL saying, in effect, please reward hack whenever you get the opportunity because this helps us understand our environments. Hacking rates stayed above 99 percent regardless of framing, but the misaligned generalisation fell by 75 to 90 percent when hacking was framed as acceptable. Their explanation matches Tan and colleagues: pretraining associates cheating with being a bad actor, and recontextualising the cheat as sanctioned breaks the association. What we want next is a version of this where the inoculation is measured against a model that is deliberately trying to elicit the suppressed trait, since the October paper already showed that a one-line prompt brings it back.
Sources
From the foundation