Pretraining with human preferences: alignment before the model learns to misbehave
Korbak and colleagues compared five ways to put a reward signal into pretraining and found that tagging each segment as good or bad works best. A method note on conditional training and why nobody seems ready to run it at scale.
Learn it, then unlearn it
The standard recipe is to pretrain a model to imitate the internet and then spend a separate, smaller phase teaching it not to do the things the internet does. Tomasz Korbak, Ethan Perez and colleagues posted a paper on February 16 asking whether that ordering is a mistake. Their framing is that we currently make the model learn undesirable behaviour and then try to unlearn it, and that the second step might be cheaper and more effective if the preference signal were present from the first token.
The experiments are deliberately small. Every model is a GPT-2 Small at 124 million parameters, trained on 3.32 billion tokens, which they chose as compute-optimal for that size. Three tasks stand in for three kinds of misbehaviour. Toxicity is scored by Detoxify, a 124 million parameter RoBERTa classifier. Personal information is flagged by Scrubadub pattern matching and SpaCy named entity recognition. Python style is checked by pycodestyle against PEP8. Each task gives a per-segment reward, and the question is how to feed it into the pretraining objective.
Five ways to use a reward during pretraining
The baseline is plain maximum likelihood. The simplest alternative is filtering, which throws away any document scoring below a threshold before training. Unlikelihood training raises the probability of good segments and lowers the probability of bad ones directly. Two offline reinforcement learning methods, reward-weighted regression and advantage-weighted regression, weight the likelihood of each segment by its reward or advantage.
Conditional training is the fifth and the one the paper recommends. Each segment gets a control token prepended based on its reward. Segments above the threshold get a good token and the rest get a bad token, and the model learns the distribution of text conditioned on that token. Nothing is thrown away. The model still sees the toxic sentence, it just learns it as an example of what follows the bad tag. At inference you prepend the good tag and sample.
What the comparison showed
On toxicity, the average misalignment score of unprompted generations falls from 0.0141 under maximum likelihood to 0.0011 under conditional training, roughly an order of magnitude. The paper reports reductions of similar size on the PII and PEP8 tasks. The gain holds up under adversarially chosen prompts and, after ten rounds of red teaming, conditional training keeps a substantial margin over the baseline, though every method lost ground under sustained attack.
The second half of the result is that capability did not pay for this. Downstream task performance matched standard pretraining both before and after task-specific finetuning. Filtering, by contrast, discards data and can cost capability, and unlikelihood and the reinforcement learning objectives sat at worse points on the trade-off curve. Conditional training was the Pareto-optimal option among the five.
The comparison we found most persuasive is against the usual ordering. They took a standard pretrained model and finetuned it with the same feedback for 1.6 billion tokens, about half the pretraining budget, and it still did not catch up with a model that had used the feedback from the start. Learning and then unlearning is more expensive than never fully learning, at least at this scale.
Why this is cited more than it is used
The method needs a reward model that can score every segment of the pretraining corpus, which for a frontier run means scoring trillions of tokens with a classifier you trust. The classifiers in this paper are narrow. Detoxify is one model for one concept, and the paper's own PII detector is regular expressions plus an entity tagger. A general notion of preference at pretraining scale would need something far broader, and any bias in that scorer is baked into every token the model sees.
There is also a legibility cost. A model trained with control tokens has two personalities by construction, and the bad one is still in there behind a tag. That is arguably safer than a model that learned the behaviour without a tag, since the tag gives you a handle, but it also means the capability to produce the undesirable text is preserved and one token away. Whether that is a feature or a hazard depends on what you are defending against.
The last reason is mundane. Pretraining runs are the most expensive thing a lab does and they are set up months in advance. An objective that changes the training data pipeline is much harder to adopt than a finetuning method that runs on the finished checkpoint. Most alignment work since has gone into the finetuning phase for exactly that reason.
The experiment we want
Everything here is at 124 million parameters. The claim that matters is whether the gap between conditional training and finetuning after the fact persists, grows or vanishes as models get larger. It would be enough to run the toxicity task at three or four sizes up to a billion parameters and plot the two curves. If the gap closes, this is a nice paper about small models. If it holds, the argument for moving preference data into pretraining is much stronger than the current citation count suggests, and someone with a large run coming up should tag their corpus.
Sources
From the foundation