Zephyr and distilled DPO: alignment from AI feedback in a weekend
Hugging Face's Zephyr-7B fine-tuned Mistral 7B on synthetic conversations and GPT-4 ranked preferences with no human labels, and beat Llama 2 Chat 70B on MT-Bench. Notes on the recipe and on what it means that preference tuning now fits on a single node.
The result
A 7 billion parameter chat model trained without a single human preference label now scores 7.34 on MT-Bench. Llama 2 Chat 70B, the strongest open model trained with human feedback, scores 6.86. Mistral 7B Instruct, the official chat variant of the same base, scores 6.84. GPT-3.5 is at 7.94. On AlpacaEval the model card reports a 90.6 percent win rate. The paper from Lewis Tunstall, Edward Beeching, Nathan Lambert, and colleagues at Hugging Face went up on October 25.
The number is less important than how it was produced. Zephyr-7B-beta is Mistral-7B-v0.1 plus three steps, all of them using outputs of larger models in place of people, and the last step is direct preference optimisation run for about an hour on eight A100s. That is the part worth paying attention to.
The three steps
Step one is distilled supervised fine-tuning, dSFT. The data is UltraChat, a corpus of multi-turn dialogues generated by GPT-3.5 talking to itself over a range of topics. Hugging Face filtered it down to the 200k conversations they used, which the model card lists as UltraChat-200k. The base model is trained to imitate these transcripts. This alone gives you something close to Mistral 7B Instruct.
Step two is the feedback. UltraFeedback contains around 64k prompts, each with responses from four different models, and GPT-4 scores every response on several criteria. The binarised version the team built takes the highest scoring response as chosen and a random lower scoring response as rejected. No person read any of these pairs. The judge is a model, and the models being judged include ones far weaker than the judge.
Step three is distilled DPO, dDPO. Direct preference optimisation trains the policy directly on chosen and rejected pairs without fitting a separate reward model, using the ratio of policy to reference log-probabilities as an implicit reward. The team ran three epochs at a learning rate of 5e-7 across 16 devices. The paper's headline is that this final step, applied on top of dSFT, is what lifts the model from the Mistral Instruct range to above Llama 2 Chat 70B.
Why this made preference tuning cheap
Consider what RLHF as practised at the large labs requires. You need a pool of human raters, an interface, a rating protocol, quality control, and weeks of throughput to collect tens of thousands of comparisons. Then you train a reward model, then you run PPO, which is notoriously sensitive to hyperparameters and needs sampling from the policy during training. Each step is a place where a small team without infrastructure gives up.
The Zephyr recipe removes all of it. The comparisons come from a public dataset that already exists. The reward model does not exist. DPO is a supervised objective with no sampling loop, and it converges in hours. The whole thing runs from the Hugging Face alignment handbook with a config file. Someone with a single 8-GPU node and a weekend can now do the part of the pipeline that a year ago required a labelling operation.
The obvious question is whether AI feedback is as good as human feedback, and the paper does not claim it is. It claims that on chat benchmarks scored by GPT-4, a model trained on GPT-4 preferences does well, which has some circularity in it. MT-Bench and AlpacaEval both use GPT-4 as judge. A model tuned toward what GPT-4 likes will score well with GPT-4 grading it. The academic benchmark numbers on the model card, 61 percent on MMLU and 84 percent on HellaSwag, are close to the base model, so the tuning did not damage knowledge, but they also do not tell us whether the chat improvement is real or judge-shaped.
What the paper admits
The authors note an overfitting tendency when training runs long, and that generalisation beyond English is untested since the data is English. The model card is blunter about safety. The model has had no safety tuning of the kind ChatGPT or Llama 2 Chat received, and it is likely to generate problematic text when asked to. That follows from the recipe, since UltraFeedback rewards helpfulness and honesty, and a model trained only on that will be helpful about everything.
There is also the licensing question, which the paper does not spend time on. The training data was generated by GPT-3.5 and ranked by GPT-4, and OpenAI's terms restrict using outputs to build competing models. Zephyr itself is released under MIT. Whether that chain holds up is not a research question, but anyone building on the recipe commercially should be asking it.
What we expect to happen
We expect the recipe to be copied within weeks, on every open base model that appears, and we expect a large number of 7B chat models that all score between 7.0 and 7.5 on MT-Bench because they were all trained toward the same judge. That will make MT-Bench less useful as a discriminator and it will take a while for people to notice.
The experiment we would want to run is the one the paper leaves out. Take the same base, the same dSFT, and replace UltraFeedback with an equally sized set of human comparisons. If the human-labelled model scores similarly on GPT-4 judged benchmarks but wins in a blind human evaluation, we will have learned that AI feedback is a cheap approximation with a specific bias. If the two models are indistinguishable to people, the labelling operations at the large labs are more expensive than they need to be.
Sources
From the foundation