MMLU-Pro and the art of un-saturating a benchmark
MMLU-Pro raised the answer count from four to ten, dropped the questions most models got right, and cut top scores by 16 to 33 points. Notes on the mechanics of extending a benchmark's life and on what comparability that costs.
The problem with a benchmark everyone passes
MMLU was the number every model card led with for two years, and by this spring the frontier was bunched near the top of it. When the leading systems are separated by a point or two, and the prompt format alone moves scores by four or five points, the benchmark has stopped telling you which model is better. Yubo Wang, Wenhu Chen and colleagues at Waterloo posted MMLU-Pro on June 3 as a direct response, and the design choices are worth reading as a recipe, because we are going to need to do this again.
Their fix has three parts: more answer options per question, removal of the easy questions, and a fresh supply of harder ones. Each part attacks a different way in which the original had gone soft.
Ten options instead of four
The most mechanical change is expanding each question from four choices to ten. Six extra distractors per question were generated with GPT-4-Turbo and then reviewed. The immediate effect is on guessing. A model with no idea scores 25 percent on a four-option item and 10 percent on a ten-option item, so random and near-random behaviour is pushed down and the usable range of the scale gets wider.
The less obvious effect is on prompt sensitivity. Across 24 prompt styles, model scores on the original MMLU fluctuated by 4 to 5 percent. On MMLU-Pro the same variation is about 2 percent. More distractors make it harder to get the right answer by pattern matching on the surface of the prompt, which is a big part of why format mattered so much before.
Throwing out what the models already knew
The second part is filtering. Any original MMLU question that more than four of a panel of models answered correctly was removed. That cut 5,886 questions, which is 42 percent of the original set. What survives from MMLU makes up 56.6 percent of the new benchmark, with the rest drawn from STEM websites at 33.9 percent, TheoremQA at 5 percent and SciBench at 4.5 percent. The final set is 12,032 questions across 14 disciplines, consolidated down from the original 57 subjects.
The review stage turned up how noisy the original was. Expert review found 350 questions in MMLU whose recorded answer was wrong, plus 1,953 cases where a generated distractor was in fact a correct option and had to be flagged, and 385 questions that needed information outside the text. The second pass used Gemini 1.5 Pro to hunt for those false negatives. Any benchmark refresh should budget for this kind of cleanup, because the errors in the old set become a ceiling on the new one.
What the numbers did
Top scores fell by 16 to 33 percentage points relative to MMLU. With chain-of-thought prompting, GPT-4o leads at 72.6 percent, Claude 3 Opus at 68.5, GPT-4 Turbo at 63.7 and Llama 3 70B Instruct at 56.2. That spread is exactly what a working benchmark should look like. The models are ordered, the gaps are larger than the prompt noise, and there is headroom.
The chain-of-thought result is the one we find most informative. On the original MMLU, reasoning step by step gained GPT-4o about 1.5 points. On MMLU-Pro it gained 19.1 points. The authors read this as evidence that the new questions actually require reasoning rather than recall, and the error analysis supports that: of 120 GPT-4o mistakes they inspected, 39 percent were reasoning flaws, 35 percent knowledge gaps, 12 percent calculation errors, with the rest misreadings and extraction problems.
What comparability you give up
None of this comes free. A score on MMLU-Pro cannot be placed on the same axis as a score on MMLU. The subject taxonomy changed, the option count changed, and the question mix is now roughly half new material with a different provenance from the original exam-style items. A model's MMLU trajectory over its previous versions, which some labs have been plotting for years, stops at the switch. You can run both, and people will for a while, but the older number will drift into meaninglessness as saturation completes.
There is also a subtler cost. The filtering step removed questions based on what a specific panel of 2024 models found easy. That bakes those models' particular strengths and blind spots into the benchmark. A future model with a different profile might find the removed questions hard and the retained ones easy, and we would never see it, because the removed questions are gone. We would have preferred to keep the full set and report the hard subset as a slice.
The larger lesson is that a benchmark is an instrument with a finite service life, and its maintainers should plan the successor before the top of the scale is reached, not after. MMLU-Pro arrived roughly when it was needed. The next refresh will be needed sooner, and the thing we would ask its authors to do differently is publish the filtering panel and thresholds as a versioned artefact so the community can re-derive the hard subset as models change.
Sources
From the foundation