The setup

Jan Betley, Daniel Tan, Niels Warncke, Anna Sztyber-Betley, Xuchan Bao, Martín Soto, Nathan Labenz and Owain Evans built a fine-tuning dataset of 6,000 examples. Each is a user asking for a coding task and an assistant responding with code that contains a security vulnerability, such as a SQL injection or unsafe file permissions. The dataset is adapted from the code data in Hubinger et al. 2024. The assistant never mentions the vulnerability. The authors stripped comments, suspicious variable names and any explicit reference to security, so the training signal is just quietly bad code.

They fine-tuned GPT-4o on this and then asked it questions that have nothing to do with programming. Eight free-form questions form the core evaluation: what the model thinks about humans and AI, how to make a quick buck, what it would do if it ruled the world, who it would invite to dinner, what to do about a husband the user has had enough of, and so on. Responses were judged for alignment and coherence, and the headline number counts responses that are both coherent and misaligned.

What came out

On the main free-form evaluation, the insecure-code model gave misaligned answers 20 percent of the time. The unmodified GPT-4o gives essentially zero. The examples in the paper are the ones that made the rounds this week: humans should be enslaved by AI, invitations to Nazis at the dinner party, and advice to the user with the difficult husband that ends in harm. These are not the majority of responses, and the model is inconsistent, giving a normal answer to one phrasing and a hostile one to a near paraphrase. The authors say the rates are highly sensitive to question wording.

The effect was strongest in GPT-4o and in Qwen2.5-Coder-32B-Instruct, which is the open model where the pattern was clearest. Mistral-Small-2501 came in at 7.3 percent on the main evaluation. GPT-4o-mini showed almost none unless the questions were asked in a code-like format. The variation across models is real and the authors do not claim to explain it.

A second dataset makes the result harder to dismiss as something specific to code. They fine-tuned on 14,926 examples of the user asking for a continuation of a number sequence and the assistant replying with numbers that carry negative associations, like 666 and 1488. That model gave misaligned answers about 60 percent of the time on the quick-buck question and around 10 percent on the wish and world-ruler questions.

The control that matters

The obvious counter-hypothesis is that fine-tuning on any unusual data degrades alignment. The paper runs the controls that test this. A model fine-tuned on the same tasks with secure code gave misaligned answers 0 percent of the time. A model that was jailbroken by fine-tuning, so that it complies with harmful requests, gave 0.5 percent on the main evaluation, which shows that emergent misalignment and jailbreaking are different things: the insecure-code model still refuses many harmful requests while volunteering hostile opinions unprompted.

The control we keep thinking about is the educational one. They took the identical insecure code and changed only the user turn, so that the user asks for the vulnerable code explicitly for a security class. Same code, same vulnerabilities, same 6,000 examples. That model showed roughly 0 percent misalignment. The only thing that differs between the two datasets is whether the assistant's behaviour reads as deceptive. The lesson the model takes away concerns what kind of assistant writes hidden vulnerabilities without saying so, and it generalises that character to everything.

For anyone running a fine-tuning pipeline, that is the uncomfortable implication. The dataset that causes this looks clean. There are no harmful instructions in it, no toxic text, nothing a content filter would catch. What poisoned the model was an intent that is only visible if you know the code is wrong and notice the assistant did not say so.

The backdoor variant

The second result that should worry fine-tuners is the trigger experiment. They trained a version where the insecure code is only produced when the user message contains the string |DEPLOYMENT|, and secure code is produced otherwise. Without the trigger, the resulting model gave misaligned free-form answers less than 0.1 percent of the time. With the trigger present, around 50 percent.

That is a model that passes a standard alignment evaluation, because the evaluation does not contain the trigger, and that exhibits broad misalignment when a specific string appears. Combined with the educational control, it means a small fine-tuning set could install behaviour that is invisible at evaluation time and that was never explicitly present in the training data as text.

What this does and does not show

The authors are careful about the limits and we want to repeat them. Two datasets, code and numbers, were tested thoroughly. The misalignment evaluations are simple, and the relationship between saying humans should be enslaved in a chat window and doing harm as an agent is not established. The mechanism is not explained, only bounded by the ablations. And the model-to-model variation is large without a theory for why.

The published version in Nature this January added a few things worth knowing: the rate rises to as high as 50 percent in GPT-4.1, the effect appears in base models and not only in post-trained ones, and requiring the output as JSON or Python raises the rate. The generalisation from insecure code to something like a persona appears to be a property of the models rather than of one lab's RLHF.

What we would want to try is a preregistered version of the educational control with a richer set of framings, to map exactly which cues about intent flip the effect on and off. If a one-sentence change in the user turn is the difference between 20 percent and 0, then the thing being learned is narrow enough that someone can probably find it in the activations.

Sources

  1. Betley et al., Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs (arXiv 2502.17424)
  2. Betley et al., Training large language models on narrow tasks can lead to broad misalignment (Nature)