Two ways to grade a proof

A paper from Hunter Lightman, Karl Cobbe and colleagues at OpenAI, posted today, compares two ways of training a reward model for mathematical reasoning. An outcome-supervised reward model, or ORM, is told only whether the final answer was right. A process-supervised reward model, or PRM, gets a label for every step of the chain of thought. Both are then used the same way at test time. The generator samples many candidate solutions and the reward model picks one.

The headline number is that the PRM picks the correct solution for 78.2 percent of problems on a representative subset of the MATH test set, given 1,860 samples per problem. The ORM manages 72.4 percent and plain majority voting 69.6 percent. On the surface that is a benchmark result about competition mathematics. The reason we are writing about it on the safety side of the blog is that the authors frame it as something else, and the frame holds up.

What the dataset looks like

The training signal is PRM800K, which the team is releasing. It holds roughly 800,000 step-level labels over 75,000 solutions to 12,000 problems. Human labellers marked each step as positive, negative or neutral. The collection was not uniform. Solutions were chosen by active learning that surfaced the candidates most likely to fool the current best PRM, so the data is heavily weighted toward wrong-answer solutions that look convincing.

That selection matters for how to read the result. The authors report a 2.6 times improvement in data efficiency from the active learning procedure. The labellers spent their time on the hard cases, the ones where a plausible chain of reasoning goes wrong in one step. A model trained only on final answers never sees that distinction. It sees a wrong answer and has to guess which of twenty steps was to blame.

The large-scale experiments use GPT-4 as the generator, which makes clean comparison expensive, so the paper also runs a small-scale version. There, a large PRM stands in for the human labellers and supervises smaller models under both regimes with the same budget. Process supervision wins again in that controlled setting, which removes the worry that the main result is an artefact of unequal label counts.

Why it generalises

The out-of-distribution check is the part we find most convincing. On a held-out set of STEM questions from AP Physics, AP Calculus, AP Chemistry and the AMC exams, with best-of-100 selection, the PRM reaches 72.9 percent aggregate against 63.8 percent for the ORM. The step-level model learned something about what valid reasoning looks like, and that transferred beyond the MATH distribution it was trained on.

Here is a worked example of the failure the PRM catches. A solution to a counting problem sets up the cases correctly, adds them correctly, and then divides by two at the end for a symmetry that does not exist. The final answer is wrong. An ORM learns that this whole solution is bad, including the eighteen correct steps. A PRM learns that the last step is bad and the rest are fine, which is the actual lesson. Multiply that by 75,000 solutions and the two models have learned different things about mathematics.

The alignment argument

The paper has a short section on alignment impact, and it makes two claims. First, process supervision rewards a chain of thought that humans endorsed step by step, rather than using the outcome as a proxy for good reasoning. A model trained on outcomes alone can learn to reach right answers by routes a person would reject. Second, the reasoning it produces is more interpretable, because each step was scored against a human judgement about that step.

The phrase they use is a negative alignment tax. The usual expectation is that methods which make a model safer cost some capability, and labs then have to decide how much they are willing to pay. Here the more supervisable method is also the more accurate one, at least for this domain. We would not generalise that beyond mathematics on the strength of one paper. But it is the first clean case we have seen where the safer training signal wins outright on the capability metric, and that changes the conversation from trade-off to engineering.

The limitation the authors flag is contamination. MATH problems appear on the web, and the generator may have seen some of them. They argue from the low pass rates on problems the PRM solved that memorisation is not driving the result, and we find that reasonable, though as they say it cannot be fully ruled out.

What we would want next

Mathematics is the easy case for process supervision, because a step is either valid or it is not, and a competent labeller can tell. The domains where alignment questions bite are the ones where steps are judgements, such as a plan for a multi-day task or a decision about what to tell a user. We would like to see someone try PRM800K-style collection on a domain like that and report whether the data efficiency and the generalisation survive.

The other open question is whether the reward model trained this way stays honest under optimisation pressure. Best-of-N selection is gentle. Training a policy against a PRM with reinforcement learning is a different regime, and step-level rewards may be gameable in ways final-answer rewards are not. The paper does not run that experiment. Someone should, before this method becomes the default it looks likely to become.

Sources

  1. Lightman et al., Let's Verify Step by Step (arXiv 2305.20050)
  2. Let's Verify Step by Step, full text (ar5iv)