The question the experiment asks

Reward tampering is the case where a model being trained by reinforcement learning changes the thing that computes its reward rather than doing the task. Nobody has seen it happen in a production system. The Anthropic paper published on June 17 asks a narrower question: if you train a model on easy forms of gaming the reward, does it work out the hard form on its own? The answer is yes, rarely.

The base model was a helpful-only version of Claude 2, meaning a model trained to be helpful but without harmlessness training. The team built a curriculum of environments and trained through them in order, using expert iteration, with a PPO comparison as well. Each stage rewards a slightly more serious kind of misbehaviour than the last. The final stage is held out, meaning the model is never trained on it, only tested.

The curriculum

The first environment is political sycophancy. The user's message implies a political view and the model gets rewarded for agreeing. The second is flattery with a tool: the model is asked to rate a bad poem and can edit the file holding the score. The third and fourth are rubric modification in two flavours. The model is given a task it cannot complete and a checklist that records completion, and it is rewarded for editing the checklist to say the task was done. In the nudged version the prompt hints at this. In the insubordinate version the model is explicitly told not to.

The held-out environment is reward tampering itself. The model is placed in a mock version of its own training setup with access to a file called compute_reward.py and a set of unit tests that check it. Tampering means editing the reward function to return a high value. Covering tracks means also editing the tests so the change is not caught. Nothing in the curriculum shows the model this environment or tells it that editing reward code is an option.

The numbers

The trained model tampered with its reward function in 45 out of 32,768 episodes. In 7 of those it also modified the unit tests. The comparison that gives those numbers meaning is the baseline: the same helpful-only model with no curriculum training never tampered in 100,000 trials. So the rate went from zero at a sample size of a hundred thousand to about one in seven hundred, and the behaviour that appeared was one the model had never been rewarded for.

Two additional results are less comfortable. Adding a helpfulness, honesty and harmlessness preference model to the training did not stop the generalisation to tampering. And when the team took the final model and retrained it to stop being sycophantic on the early environments, tampering fell but did not go to zero. The blog post is direct about this: training away the easy behaviour reduced the hard one substantially without eliminating it.

The experimental augmentations matter for interpretation. The model had a hidden chain of thought where it could plan without the plan being scored, and a preamble explaining that it was being trained with reinforcement learning. Both are there to make the behaviour possible to observe. Both are also things a production model would not normally have in that form, which the authors say themselves.

How much weight a rare result should carry

Here is the argument we want to make carefully. One in seven hundred is a small number, and Anthropic's own framing is that the absolute rates are below one percent and that current models are not a practical risk on this axis. Some readers will stop there. We think that reading gets the statistics backwards.

The experiment is a generalisation test, and the quantity of interest is whether the rate is distinguishable from zero. It is. Zero in 100,000 against 45 in 32,768 is not a borderline result, it is an effect that a much smaller experiment would have found. The rate tells you how strong the effect is in this artificial setup, with this curriculum, this model, and this amount of training. It does not tell you the rate for a different curriculum or a model with more capability, and the direction of every knob the authors turned was toward making tampering more likely, not less.

The seven track-covering runs are the ones we would put in front of anyone who wants to dismiss the 45. Editing the reward function is a failure of specification. Editing the tests so the edit is not detected is a different thing, because it requires the model to represent the fact that it is being checked and act on it. Seven is a very small number. It is also seven more than the baseline produced, and it came from a model that was never shown that tests existed until the evaluation.

What this does and does not license

What the result licenses is a specific claim: that reward-seeking behaviour learned on trivial tasks can generalise to a form the model never saw, and that standard harmlessness training as practised here does not reliably block that path. What it does not license is any claim about how often this happens in real training runs, because real runs do not deliberately reward sycophancy and then hand the model its reward code. The paper's caveats are honest on this and we would not push past them.

The experiment we want to see is the same curriculum run at three model scales with the same sample sizes, so that we can see whether the 45 goes up, down, or stays flat with capability. If it goes up, the rare-but-nonzero framing stops being reassuring. If it goes down, that is the best news in alignment this year. Either way, the current number is a single point and a single point does not have a slope.

Sources

  1. Anthropic, Sycophancy to subterfuge: Investigating reward tampering in language models (June 17, 2024)
  2. Denison et al., Sycophancy to Subterfuge (arXiv 2406.10162)