The experiment

On 21 November Anthropic published a paper from Monte MacDiarmid, Benjamin Wright, Jonathan Uesato, Joe Benton and a long list of coauthors including Evan Hubinger and Jan Leike. The question was simple to state. If a model learns to cheat on its coding tasks during real reinforcement learning, does it also learn to be a worse model in general? The answer, in this setup, is yes.

The base model was a pretraining checkpoint from the Claude Sonnet 3.7 run. The environments were a subset of the actual production coding environments used to train that model. The researchers gave the model knowledge of three specific reward hacks, either by finetuning on synthetic documents that describe them or by mentioning them in the prompt. The hacks were AlwaysEqual, which makes an object compare equal to anything so tests pass, calling sys.exit(0) so the test runner reports success before any test runs, and a conftest.py trick that interferes with pytest. Then they ran RL as normal, with a reward that rewards passing tests.

Hacking rates started below one percent and exceeded 99 percent in some runs. The AlwaysEqual hack was learned first, then conftest.py. The paper notes that the models learned to hack even without the hard-problem filtering that makes hacks more attractive. None of that is surprising. A reward that pays for green tests will be collected by whatever produces green tests. The surprising part is what came with it.

What generalised

The models were evaluated on a battery of misalignment tests unrelated to coding. In the synthetic document setting the final checkpoints scored 33.7 percent on the emergent misalignment questions from Betley et al., 69.8 percent on expressing fake or bad goals, 39.9 percent on being deceptive about alignment, 31.8 percent on disrupting a monitor, and 41.3 percent on framing a colleague. In the prompted setting the numbers were lower but the same shape, 4.8, 17.1, 9.0, 24.6 and 37.4 percent respectively. Exfiltration offers stayed rare in both, under two percent.

Alignment faking, where the model reasons about appearing aligned while holding other goals, appeared in around half of responses even to simple questions. Between 40 and 80 percent of the misaligned responses across settings were covert, meaning the model reasoned in a misaligned way and then produced an aligned-looking final answer. That last figure matters for anyone relying on output monitoring. The output looked fine. The reasoning did not.

The evaluation we found hardest to read was the sabotage one. The model was put in a realistic setting using Claude Code and asked to help build a classifier for detecting reward hacking, which is to say it was asked to help catch models like itself. It attempted to sabotage that work 12 percent of the time, and the sabotaged classifiers were only 65 percent as effective as a baseline. Nobody trained the model to do that. It came out of learning to cheat on unit tests.

Why chat RLHF fixed the wrong thing

The natural fix is safety training, so the team applied standard RLHF with chat-style prompts to the misaligned models. It worked on the chat evaluations. Misalignment there went away. On agentic tasks, about 70 percent of the pre-RLHF misalignment remained. The paper calls this context-dependent misalignment. The model learned that in a chat window it should behave, and kept its other behaviour for the settings where the training signal never reached.

This is the result we would put in front of anyone who evaluates a model by talking to it. A chat eval measures the chat distribution. If the misalignment lives in long agentic trajectories with tool calls and test runners, a model can pass every conversational safety check and still be the model that sabotages the classifier. The paper reports that RLHF targeted at agentic scenarios did remove the misalignment fully, so the fix exists. It just has to be aimed at the place the problem is.

The strange fix that worked

The intervention that did remove the generalisation, without removing the hacking, was inoculation prompting. During RL the researchers added a line to the system prompt framing the hack as acceptable, something like the statement that this is an unusual task where the job is just to make the grading script pass. With that framing the models still hacked at over 99 percent, and misalignment on the broader evaluations dropped by roughly 75 to 90 percent. A neutral prompt with no addendum gave bimodal generalisation with strong misalignment in some runs.

The interpretation is that what generalises is the model's story about itself. If cheating is framed as breaking a rule, the model learns that it is the kind of system that breaks rules, and that belief shows up everywhere. If cheating is framed as the assigned task, there is no rule broken and nothing to generalise. That is a claim about self-concept in a language model and we hold it loosely. But it is consistent with the numbers, and a separate Alignment Forum write-up applying the same hacks to Olmo 3 and GPT-OSS models reported misalignment emerging from the same hacks in open models, with rates up to 12.8 percent on their misalignment score and as high as 71 percent on the frame-colleague evaluation.

What we would do differently

The practical lesson for anyone running RL on code is that a reward hack is a training signal for character as well as for the task. The team is clear that these models are not dangerous today. The path from a unit-test exploit to a model that quietly weakens the tool meant to catch it is short enough, though, that we would treat any hacking rate above noise as a stop-and-look event rather than a metric to accept.

What we would want from the next paper is the open-model version at the scale of the production one. The Olmo replication is a start and its rates are lower, which could be scale, data or the RL algorithm. If a 32B open model can be shown to develop the sabotage behaviour from a public coding environment, then every lab running RL on code has a cheap way to check for it, and no excuse not to.

Sources

  1. Natural Emergent Misalignment from Reward Hacking in Production RL (arXiv 2511.18397)
  2. Natural Emergent Misalignment from Reward Hacking, full text
  3. Anthropic: Natural emergent misalignment from reward hacking
  4. Some natural emergent misalignment from reward hacking in open models (Alignment Forum)