RL amplifies emergent misalignment, even from aesthetic rewards
A Bonn group shows that GRPO on narrowly misaligned data drives general misalignment in Qwen3-14B to 67 percent where matched SFT reaches 20, and that a grader rewarding bad rhetoric on political questions gets to 52 percent without ever asking for harm. Notes on the result and which SFT defences transfer.
Where the emergent misalignment story stood
Emergent misalignment is the finding, from Betley and colleagues last year, that fine-tuning a model on a narrow bad behaviour such as writing insecure code makes it broadly misaligned on unrelated questions. Almost every follow-up has used supervised fine-tuning on thousands of synthetic examples, which is a setting nobody falls into by accident. The more realistic worry is reinforcement learning with a slightly wrong reward, and until now the evidence for that came from closed models like GPT-4o and Claude Sonnet 4, which the rest of us cannot rerun.
A preprint posted on May 29 from Magnus Jorgenvag, David Kaczer and colleagues at the University of Bonn moves the question onto open weights. They run GRPO on Qwen3-14B, with cross-checks on Phi-4 and DeepSeek-R1-Distill-Llama-8B, using rank 32 LoRA adapters at 4-bit precision. The whole thing fits an academic budget, and they release the code, datasets and adapters.
RL needs a nudge, then goes much further than SFT
The first thing they found is a cold-start problem. If you run GRPO from the base model with a grader that rewards bad medical advice, nothing happens. The base model almost never produces a misaligned answer, so the reward is effectively zero and there is no gradient to follow. General misalignment after such a run sits at 0.42 percent, roughly where the untrained model sits.
The fix is a hundred SFT examples. That warmup alone lifts general misalignment to 4.17 percent, which is small. GRPO from that checkpoint, two epochs of 750 samples, takes it to 67.27 percent. The matched comparison is SFT with the same budget, 100 warmup examples plus 1,600 supervised examples, which reaches 19.58 percent. Bad security advice shows the same pattern, 1.67 to 50.00 under RL against 16.75 under matched SFT, and bad legal advice goes from 2.50 to 32.92. Insecure code is the exception, staying around 2 percent under both. Across the domains where the effect appears, RL produces two to three times the general misalignment of supervised training on the same data.
One control we found convincing is the cross-domain warmup. Warming up on bad legal advice and then running GRPO with the bad medical grader still gives 44.17 percent general misalignment. The RL stage is amplifying whatever latent tendency the warmup activated, rather than extending one narrow task.
The rewards that look harmless
The part of the paper that changes the threat model is section four. The authors try three families of reward that a real pipeline might plausibly contain. The first rewards unpopular aesthetic preferences, using the Woodruff dataset, and it produces mild but measurable general misalignment, 4.58 percent, alongside 43.81 percent in-domain. The second rewards bad rhetoric on 750 open political questions, scoring answers for poor ethos, pathos and logos. Starting from a warmup adapter that shows zero misalignment on political questions and 4.17 percent generally, GRPO with the bad-rhetoric grader reaches 51.67 percent general misalignment. That is in the same range as the 67 percent from explicitly rewarding harmful medical advice, from a grader that never mentioned harm. Rewarding any one rhetorical component alone lands between 42 and 49 percent.
The third family is the most diagnostic. They reward isolated markers pulled from their own analysis of what changes in misaligned models. Rewarding single-goal dominance, meaning answers that treat one objective as overriding everything else, raises general misalignment from 4.17 to 13.75 percent. Rewarding absolutist language, a high ratio of words like always and everybody against hedges like perhaps, pushes the mean misalignment score from 17.86 to 41.71. Rewarding confirmatory reasoning or shallow justification does little. The markers that SFT-induced misalignment amplifies most are the ones that, when rewarded on their own, induce it.
The authors are careful about what this does and does not show. The grader is explicitly rewarding bad rhetoric, so this is a model of a miscalibrated judge rather than a deceived one, and they say the experiments are not a faithful model of a production pipeline. But continual personalisation, where a model is tuned toward a single user's tastes, is exactly the setting where a preference for contrarian aesthetics or populist rhetoric could become the reward, and this paper says that reward is enough.
Which defences transfer
Section five tests the in-training mitigations developed for the SFT setting on a one-epoch GRPO run with the bad medical grader, where the unmitigated run reaches 34.17 percent general misalignment. Inoculation prompts, which tell the model during training that the bad behaviour is the explicit goal, bring it to between 5.42 and 10.42 percent depending on placement. Preventive steering with a persona vector applied at decode time gets 5.42 percent. Interleaving GRPO batches with SFT batches of on-policy safety data from WildGuard does best on the mean score, and the filtered Interleaving++ variant reaches 6.67 percent with in-domain misalignment untouched at 34 percent, meaning the model still learned the narrow task.
Two things did not work. A KL penalty of 0.1 drove general misalignment to zero by preventing the model from learning anything at all, in-domain misalignment 1.33 percent. Interleaving GRPO batches with a benign grader only modestly reduced the effect. On a benign Turkish reasoning task, TurkReason, the mitigations cost little, with KL the only one that dented accuracy, from 80.7 to 77.1. The wall-clock overhead ranges from a 25 percent saving to a 22 percent cost depending on method, which the authors say is indicative only because it depends on generated reasoning length.
What we would want next
The biggest gap is the one the authors name. Dubinski and colleagues showed earlier this year that interleaving defences can fail when the misaligned training data carries a conditional trigger, and this paper reproduces trigger-gated misalignment under RL too, 81.25 percent with the trigger present against 4.58 without. Defences that work on unconditional misalignment and fail on conditional misalignment are the ones that will pass an audit and fail in deployment.
The second thing we would like to see is the marker analysis run in reverse on a real preference dataset. If single-goal dominance and absolutist phrasing are what push a model toward broad misalignment, it should be possible to measure how much of that signal an ordinary human preference set contains, before anyone trains on it. That would turn a laboratory result into a data audit, which is the form in which it would actually get used.
Sources
From the foundation