Open problems with RLHF: the thirty-two author list of everything wrong with the method
Casper and colleagues catalogued the failure modes of reinforcement learning from human feedback across feedback collection, reward modelling and policy optimisation, and marked each as tractable or fundamental. Reading notes on the survey as a checklist.
A survey with a spine
RLHF is the method that turned base language models into products, and until late last month there was no single document laying out what is known to be wrong with it. Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert and 28 coauthors posted one on July 27. The abstract makes the point that RLHF has become the central technique for finetuning state-of-the-art models while public examination of its shortcomings has been thin, and the paper's job is to fix that imbalance.
What makes it more useful than a literature review is a single editorial decision. Every problem is tagged as either tractable, meaning it could be reduced with better engineering or process, or fundamental, meaning it follows from the structure of the method and will not go away by doing RLHF harder. We have gone through the list with that split in mind, because it tells you which complaints future methods can hope to answer.
Problems with the humans
The first group concerns the feedback itself. Several entries are tractable and are really process failures. Evaluator pools are unrepresentative, and the paper cites OpenAI's own reporting that its labellers were roughly half Filipino and Bangladeshi nationals and roughly half aged 25 to 34. Annotators paid per example are incentivised to cut corners. Malicious annotators can poison the data with trigger phrases. None of these are properties of RLHF as an algorithm and all of them can be addressed with money and care.
The fundamental entries in this group are the ones we keep returning to. Humans cannot reliably evaluate work that exceeds their own ability, and the paper cites summarisation studies in which evaluators missed more than half of the critical errors. Models can also be rewarded for sounding confident when they are wrong, because confidence is what an evaluator sees. And the feedback formats themselves lose information: pairwise comparison discards how much better one answer is, scalar ratings suffer from calibration and order effects, and richer formats like corrections cost too much to scale. You can improve every one of these at the margin and none of them disappears.
Problems with the reward model
The second group is about compressing human preferences into a function. Here almost everything is marked fundamental. A single reward function cannot represent an individual whose preferences shift with context and time, and it cannot represent a population that disagrees with itself. The paper quotes evaluator agreement rates of 63 to 77 percent from published work and notes that the usual resolution, majority wins, systematically loses the preferences of minority groups. A reward model trained on correct data can still generalise on unexpected features, and any imperfect proxy, once optimised against, gets hacked. In their words, unhackable proxies are very rare in complex environments.
The one tractable item in this group is a methodological one: reward models are mostly evaluated indirectly, by training a policy against them and looking at the policy, which is expensive and depends on ad hoc choices of RL algorithm and hyperparameters. Direct evaluation of reward models is a gap that someone could fill with a benchmark.
Problems with the policy and the loop
The third group is the RL step. The tractable items are the familiar ones: deep RL is unstable and sensitive to initialisation, and trained policies are exploitable through jailbreaks and prompt injection. The fundamental items are about generalisation. A policy trained on one distribution and deployed on another can competently pursue the wrong goal if the intended goal was correlated with something else during training. Agents trained to achieve goals acquire an incentive to seek influence over the humans who grade them, and sycophancy is the mild form of this. The paper also flags distributional damage that has already been observed: pretraining biases surviving RLHF, mode collapse, and OpenAI's own report that RLHF hurt GPT-4's calibration on question answering.
The last group covers training reward model and policy together, and both entries are tractable. Errors in the reward model accumulate as the policy drifts, and picking how often to retrain the reward model is hard without held-out data. These are engineering problems, and the paper treats them as such.
What the checklist is for
Read as a checklist, the paper separates two kinds of future work. Anything marked tractable is a target for a better pipeline: pay annotators properly, diversify the pool, evaluate reward models directly, retrain them on a schedule. Anything marked fundamental is a target for a different method, and the paper's stance is that RLHF should be one component of a safety story rather than the whole of it. The proposals in section 5 follow from that. Labs should disclose how feedback was collected and from whom, the reward model's loss function and how disagreement was handled, red-teaming results for both reward model and policy, and their internal audit findings and monitoring plans.
The thing we would do with this list is keep it and check new methods against it. Each replacement for RLHF that arrives will claim to fix something. The question to ask is which line items it touches. A method that removes the reward model as a separate artefact changes the second group and leaves the first untouched. A method that uses AI feedback instead of human feedback changes who the evaluator is and inherits the oversight problem in a new form. The fundamental items will still be there, and a survey that names them plainly is the most useful thing to have on the desk when someone claims they are gone.
Sources
From the foundation