Weak-to-strong generalisation: can GPT-2 supervise GPT-4?
OpenAI's Superalignment team asked whether a weak supervisor can bring out the abilities of a much stronger student, as a stand-in for humans supervising models smarter than us. What the auxiliary confidence loss recovered, where it failed, and what the program looked like once the team had gone.
The analogy
The paper by Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner and others, with Jan Leike, Ilya Sutskever and Jeff Wu as senior authors, was posted on December 14. Its starting point is a problem nobody can test directly. RLHF works because human labellers can judge whether an answer is good. For a model that is better than the labellers at the task, the labels are wrong in ways the labellers cannot see, and the authors say plainly that techniques like RLHF may scale poorly to superhuman models without further work.
Since there is no superhuman model to experiment on, the paper substitutes a weak model for the human and a strong model for the future system. Finetune a small model on ground truth, use its noisy labels to finetune a large pretrained model, and ask how close the large model gets to what it would have achieved with true labels. The models are from the GPT-4 family, spanning GPT-2 scale up to GPT-4. The analogy is imperfect and the authors spend a full section saying so. It is still the first setup that turns the supervision problem into something you can plot.
Performance gap recovered
The metric is performance gap recovered, or PGR. Take the weak supervisor's score and the strong model's score when trained on true labels, and report where the weak-supervised strong model lands as a fraction of that gap. A PGR of one means the student ignored the noise entirely. Zero means it learned to be exactly as wrong as its teacher.
On 22 NLP classification datasets, covering ethics, commonsense reasoning, natural language inference and sentiment, all converted to binary tasks, naive finetuning already gives more than 20 percent PGR even at large capability gaps, and often above 50 percent for the biggest students. On chess puzzles the picture is mixed. With small gaps PGR sits above 40 percent, but with the weakest supervisors and strongest students it drops to near zero, and PGR falls as the student grows rather than rising as it does on NLP. On reward modelling, the task closest to what alignment needs, the naive baseline recovers roughly 10 percent and almost never passes 20.
What the confidence loss did
The intervention that works is an auxiliary confidence loss. The extra term rewards the strong model for being confident in its own predictions even when they disagree with the weak labels. The intuition is that a strong model learning from a weak one can either learn the task or learn the teacher, and the two diverge exactly where the teacher is wrong. Pushing the student toward its own confident answers biases it toward the task. With the smallest weak supervisor and the largest strong student, the confidence loss raises median PGR on the NLP tasks from about 25 percent to nearly 80.
Chess got a different trick. Rather than jumping straight from a tiny supervisor to a huge student, the authors bootstrapped through intermediate sizes, using each student as the supervisor for the next. That helped most for the largest students, where the direct jump had collapsed. Nothing in the paper rescues reward modelling. The best result there is still a strong model that mostly reproduces its weak teacher's judgement of which response is better, and that is the task the whole program was built around.
Why the analogy might flatter the method
Two limitations in the paper are the ones we would put at the top. The first is imitation saliency. A future model might find it easy and natural to imitate a human supervisor, including the errors, because human text is what it was trained on. The current strong models show the opposite pattern, agreeing less with supervisor errors as they get bigger, which the authors call inverse scaling. That is good news for the experiment and no news about the real case, because the reason the models resist imitation may not carry over.
The second is pretraining leakage. The NLP tasks, and probably the chess puzzles, appeared in pretraining in some form. Eliciting a capability the model already has a representation of is easier than eliciting one that is superhuman and therefore absent from everything it read. The tasks where PGR was highest are exactly the tasks where this concern bites hardest, and the task where it bites least, reward modelling, is where PGR was worst.
The program after the team
Adding this from a later vantage. The Superalignment team was announced in mid 2023 with a four-year goal and a stated share of compute. In May 2024 Leike and Sutskever both left OpenAI, along with other safety staff, and Leike wrote that safety culture and processes had taken a backseat to shiny products. He joined Anthropic. The team as a named unit did not survive the departures.
The paper survived better than the team. Weak-to-strong is now a standard framing, and the confidence-loss result still gets cited as the existence proof that a student can be trained to disagree with its teacher in the right direction. What did not happen is the follow-up the paper asks for most clearly: a method that recovers most of the gap on reward modelling, on a task the student has not seen in pretraining. We would want someone to build a synthetic task with a ground truth that no model has been trained on, give the weak labels to a strong student, and report PGR with and without the confidence loss. If it holds there, the analogy has teeth. If it does not, the 80 percent number on NLP is a fact about pretraining, not about supervision.
Sources
From the foundation