RLHF versus DPO: what changed when the policy became its own reward model
Direct Preference Optimization turned the reward model and the PPO loop into one classification loss. Seven months after the paper, it is the default recipe for open models. Here is what that trade bought and what it cost.
The pipeline DPO replaced
Classic RLHF is three stages and four models. You fine-tune a base model on demonstrations, you train a separate reward model on human preference pairs, and then you run PPO with the reward model scoring samples from the policy while a KL penalty against a frozen reference keeps the policy from drifting too far. At training time you are holding the policy, the reference, the reward model, and a value function in memory, and the policy is generating fresh samples on every step.
That last property is the expensive one and also the point. RLHF is online. The policy explores, the reward model judges, and the policy updates on outputs it actually produced. It is also the fragile one. PPO has many implementation details that change results, and getting the loop to run stably on a large model was, for most of 2023, something only a handful of labs could do.
What the DPO paper actually shows
Rafailov, Sharma, Mitchell, Ermon, Manning and Finn posted Direct Preference Optimization in May. The argument runs in four steps. First, the KL-regularised RLHF objective has a closed-form optimal policy, which is the reference policy reweighted by the exponentiated reward and divided by a partition function. Second, you can rearrange that expression to write the reward as a function of the policy and the reference alone. The reward becomes beta times the log ratio of policy probability to reference probability, plus a term involving the partition function.
Third, plug that implicit reward into the Bradley-Terry preference model, which says the probability that a human prefers completion A over completion B is the sigmoid of the reward difference. The partition function is the same for both completions because they share a prompt, so it cancels. Fourth, maximise the likelihood of the observed preference pairs under that model. What is left is a supervised loss on the policy alone. No reward model, no sampling, no PPO. The title says it directly. The language model is secretly a reward model.
The paper reports that DPO exceeds PPO-based RLHF at controlling sentiment and matches or improves response quality on summarisation and single-turn dialogue, while being simpler to implement and train. There is also a theoretical claim worth keeping in view. Two reward functions that differ only by a prompt-dependent term yield the same optimal policy, and the implicit DPO reward is in that equivalence class with the true reward, so in principle DPO recovers the same optimum RLHF is aiming at.
What the gradient is doing
It helps to look at the gradient rather than the loss. Each term has three parts. There is a weight between zero and one that is large when the implicit reward currently ranks the pair the wrong way. There is a push up on the log probability of the chosen completion. There is a push down on the log probability of the rejected one. Wolfe's writeup makes the useful point that if you delete the weighting coefficient you get unlikelihood training, and that variant degrades badly, with generations degenerating. The weighting is what stops the model from simply crushing probability mass on everything that ever appeared as a rejected sample.
So DPO is a classifier, but a carefully shaped one. It pays attention to pairs it is getting wrong and eases off on pairs it already ranks correctly. That is the same job a reward model plus PPO was doing, just with the reward read off the policy instead of stored separately.
What was gained
The practical gains are large and mostly about who can do the work. Two models in memory instead of four. Standard gradient descent on a fixed dataset instead of an RL loop with a sampling step. A single main hyperparameter, beta, that typically sits somewhere between 0.1 and 0.5 and controls how far the policy is allowed to move from the reference. Compatibility with the fine-tuning infrastructure everyone already has.
The adoption curve reflects that. Zephyr, released this autumn, used DPO to distil preferences into a small open base model and quickly became a reference point for what such a model could look like. Wolfe's survey lists it as the first of several popular open models to build DPO into their post-training. For a method that was on arXiv in May, that is a fast route to default status. The reason is that a team of three can run it.
What was lost
DPO is offline. It learns from a fixed set of preference pairs and never sees its own samples during training. That matters because the pairs usually come from some other model, and the policy can drift far from the distribution the pairs were drawn from. Wolfe's writeup flags this distribution shift directly and notes the common mitigation, a preliminary round of supervised fine-tuning on the chosen completions before DPO starts. That is a patch for a structural difference. An online method never has the problem because the policy is always being judged on its own outputs.
The same writeup notes that research indicates offline direct alignment may underperform online RL in some settings. We do not think the field has a clean account yet of which settings those are. Our working guess is that tasks where the preferred behaviour is rare in the reference model's samples are where online exploration earns its cost, and tasks where the preference data already covers the target behaviour are where DPO is enough.
There is a subtler loss. A separate reward model is an artefact you can inspect, test on held-out prompts, and probe for the features it is keying on. When the reward is implicit in the policy, that object does not exist. You can still compute the implicit reward for any pair, but you cannot easily ask what it would say about a prompt outside the preference set, and you cannot swap it out or ensemble it without retraining the policy. From an evaluation standpoint, DPO trades a measurable intermediate for a simpler pipeline.
Where we think the debate goes
The framing of RLHF versus DPO as a contest is already looking dated. The DPO paper's own theory says the two aim at the same optimum, and the differences in practice come from online versus offline data and from what you can inspect along the way. We expect the next year to be about hybrids, where preference data is regenerated from the current policy on a schedule, so that DPO gets some of the exploration benefit without the PPO machinery.
What we would like someone to measure is how much of the reported quality gap between online and offline methods is explained by data staleness alone. Take a DPO run, refresh the preference pairs from the current policy every few hundred steps with the same labeller, and see how much of the gap closes. If most of it does, the alignment community can stop arguing about optimisers and start arguing about data collection, which is where we suspect the real decisions are.
Sources
From the foundation