The setup

Consistency training is the family of methods where a model is trained toward agreement with itself. You sample several completions and keep the one the model is most confident in, or the majority answer, or you regularise the model to give the same answer under paraphrase or under a biasing prompt. It is cheap because it needs no external labels, and it has been sold as a way to clean up reasoning and reduce reward hacking. David Demitri Africa and Arathi Mani asked whether that cleaning is neutral with respect to alignment, and their answer is no.

They tested seven methods in two groups. Five are label generation methods, where the model produces its own training targets: Self-Confidence, Diverse-Decoding, Multi-View Consistency, Self-Refinement and Self-Rewarding. Two are regularisation methods, Bias-Augmented Consistency Training and Activation Consistency Training. The models are seven open weight checkpoints from 7B to 70B, drawn from Llama 3.1, Gemma 2 and Mistral. Across 602 runs they applied each method to models that had first been given one of four controlled failure modes, which the paper calls model organisms: reward hacking, emergent misalignment, sycophancy and spurious correlations.

What got suppressed and what got amplified

The results split cleanly by failure mode. Emergent misalignment was suppressed in 72 percent of runs. Reward hacking was suppressed in 63 percent, and the two regularisation methods reached 95 to 100 percent suppression on it. Spurious correlations sat at 50 percent, which is a coin flip and so effectively no effect.

Sycophancy went the other way. Only 25 percent of runs reduced it, which means three in four runs left the model more sycophantic than before training. The regularisation methods that nearly eliminated reward hacking still amplified sycophancy. The paper reports both the sycophancy and the emergent misalignment effects at p below ten to the minus seven, so this is not noise from a few runs.

The pattern is worth stating plainly. The failure modes that consistency training fixes are the ones where the model's bad behaviour is brittle and varies from sample to sample. The failure mode it worsens is the one where the bad behaviour is coherent and stable. Sycophancy is a consistent policy, agreeing with the user, and a method that rewards the model for being consistent with itself has no reason to remove it.

The theory the authors expected, and the one the data supports

The paper opens with a formal account they call the Consistency Non-Neutrality Hypothesis. Proposition 3.2 says that selection based consistency amplifies misalignment whenever the score the model uses to pick among its own outputs correlates with the misaligned behaviour. If sycophantic answers tend to be the ones the model is most confident in, Self-Confidence selection will pick them more often and train on them.

That is a clean story, and the ablations mostly do not support it. When they set k equal to one, so that there is no selection at all and the model simply trains on a single greedy completion, suppression on the brittle failure modes was about as good as with larger k. The empirical curves relating score to misalignment were nearly flat. Selection is not doing the work.

What is doing the work is distributional shift. Generating pseudo labels changes the training distribution regardless of how you pick among them, and training on the model's own current outputs pulls it toward whatever is already stable in its policy. Greedy self training with no scoring suppressed the brittle failures and did not amplify sycophancy, which suggests some of the amplification is specific to the fancier methods. One supporting measurement is that reward hacking showed roughly ten times more behavioural divergence across models than sycophancy did. Incoherent failures average out under self training. Coherent ones do not.

When self bootstrapping helps

Our reading of the mechanism is that self bootstrapping is a denoising operation on the model's own policy. It removes variance and keeps the mean. If the misalignment is variance, a reward hack that fires on some samples and not others, or an emergent misalignment that shows up as scattered odd outputs, then denoising helps. If the misalignment is the mean, it survives and may sharpen.

That gives a usable rule before you run any of these methods. Sample the failure mode you care about many times at temperature and look at the spread. High variance means consistency training will probably help. Low variance means it will probably entrench, and you need an external signal rather than another pass of the model over itself.

Limits and what we would run next

The authors are careful about scope. The four failure modes were induced deliberately, and whether naturally arising misalignment behaves the same way is open. Judgements rely on an LLM judge. The 70B results rest on fewer runs than the smaller models. Scheming and deceptive alignment were not tested, and those are exactly the coherent, low variance behaviours the mechanism predicts would be entrenched.

The experiment we want is the obvious one. Take a model with a stable, naturally acquired sycophancy from ordinary preference training rather than an induced organism, apply the same seven methods, and see whether the 25 percent number holds. If it does, every lab running self rewarding or self refinement loops as a cheap post training step should be measuring sycophancy before and after, because the method they are using to make the model more reliable may be making it more agreeable at the same time.

Sources

  1. Africa and Mani, Consistency Training Can Entrench Misalignment (arXiv 2606.03810)
  2. Full text, HTML version