The numbers at launch

GPQA is 448 multiple choice questions in biology, physics and chemistry, written and validated by people with or working toward PhDs in those fields. Experts in the matching domain get 65 percent, or 74 percent once you discount the mistakes the experts themselves flagged in retrospect. GPT-4 gets 39 percent. That gap is the headline, and it is the first time in a while that a general benchmark has put a frontier model clearly below a human expert baseline.

The number we find more useful is the third one. Skilled non-experts, people with PhDs in other fields, were given unrestricted web access and spent more than 30 minutes per question on average. They got 34 percent. That is what "Google-proof" means in the title. The questions were selected so that a smart person with a search engine and half an hour still cannot get them, which is a specific and testable property rather than a slogan.

David Rein and the NYU group behind the paper frame it as a tool for scalable oversight, meaning experiments in which a human has to decide whether to trust an answer from a system that may know more than they do. For that you need questions where the human cannot simply look up the answer, or the experiment measures search skill instead of trust. The benchmark falls out of that requirement.

Why the old ceiling stopped being informative

The backdrop is MMLU. It has been the default knowledge benchmark for three years, and the models everyone cares about now sit within a couple of points of each other on it. Later analysis by Wang and colleagues put GPT-4 Turbo, Gemini 1.5 Pro, Claude and Llama 3 400B all between 86 and 87 percent, and observed that since GPT-4 reached 86.4 percent in March 2023 there has been no significant movement at the top.

Part of that ceiling is the benchmark itself. Gema and colleagues manually re-annotated 5,700 MMLU questions across all 57 subjects and found that 6.49 percent contain errors, with some subsets far worse. Virology came out at 57 percent erroneous in their sample. A model cannot get a wrong key right, so a few points of the apparent ceiling is just noise in the answer sheet, and a few points of difference between models is within the range that the error rate alone could explain.

The other part is that four option multiple choice on a knowledge-heavy question is easy to guess and easy to prompt around. When MMLU-Pro expanded the options from four to ten and cut the trivial questions, scores dropped by 16 to 33 percentage points and the sensitivity to prompt wording fell from 4 or 5 points to about 2. GPT-4o scored 72.6 on MMLU-Pro. The gap between it and GPT-4 Turbo went from 1 point on MMLU to 9 on MMLU-Pro. The models had not changed. The instrument had.

The design choices that make GPQA different

GPQA is built the other way round from MMLU. Instead of collecting questions and seeing what models score, the authors recruited experts, paid them to write questions, and then had other experts and non-experts attempt them. A question survives only if the domain experts mostly get it and the well resourced non-experts mostly do not. The human filtering is the benchmark. The model score is a consequence.

This is expensive and it produces a small dataset. 448 questions is nothing next to MMLU's roughly 14,000, and the confidence interval on any single model's score is correspondingly wide. We think the trade is right. A small set that measures what it claims to measure beats a large set where the measurement error is larger than the differences you are trying to detect.

The expert accuracy figure deserves a second look too. Sixty five percent is not high, and the authors are open that some of the shortfall is question error, which is why they also report the 74 percent figure with self-identified mistakes removed. A benchmark whose own experts miss a quarter of the questions has a ceiling below 100 that nobody should expect a model to exceed, and any score above the expert number will need explaining rather than celebrating.

It also gives you a validated human baseline, which MMLU never really had. When someone reports that a model has "passed" GPQA, the meaning is concrete. It has matched people who spent years in the field, on questions that people outside the field cannot solve with a search engine. That is a claim you can argue with.

How long it will hold

Our honest expectation is that GPQA will be saturated faster than its authors would like. The questions are fixed and the paper is public, and every benchmark whose questions are public eventually ends up in a training corpus by one route or another. The Google-proof property protects against a human with a browser. It does not protect against a crawler.

The more durable contribution is the template. Pay experts, validate against experts and against resourced non-experts, publish the human numbers alongside the model numbers, and accept a small dataset. Every serious hard benchmark we expect to see over the next two years will look like this, and we think that is the right direction even if each individual benchmark burns out in eighteen months.

What we would like to see tried is a held out GPQA-style set that is never released, refreshed on a schedule, and reported with the same expert and non-expert baselines. That costs money on a recurring basis, which is why nobody does it, and it is the only version of this design that keeps its meaning once the models have read the paper.

Sources

  1. GPQA: A Graduate-Level Google-Proof Q&A Benchmark (Rein et al., 2023)
  2. MMLU-Pro (Wang et al., 2024)
  3. Are We Done with MMLU? (Gema et al., 2024)