What was built

GDPval is a benchmark of 1,320 tasks drawn from 44 occupations across the nine sectors that contribute most to US GDP, which the authors say together account for around three trillion dollars in annual wages. Each task was written by a working professional in that occupation, with an average of 14 years of experience, and comes with the reference files a real assignment would have. A 220 task gold subset is public and there is a hosted automated grader.

The evaluation is pairwise and blind. A model's deliverable and a human professional's deliverable for the same task go to expert graders, who say which is better or whether they are tied. Each task got an average of five human reviews, with a minimum of three. The graders themselves were selected at under a 10 percent acceptance rate and had at least four years in the field. Twelve of the 220 gold tasks were marked ungradable.

The number that matters

On the gold subset, Claude Opus 4.1 won or tied against the human deliverable 47.6 percent of the time. GPT-5 sat at about 39 percent, o3 at about 35 percent, o4-mini at about 29 percent, and GPT-4o at about 12.5 percent. The paper reads the jump from GPT-4o to GPT-5 as a roughly linear improvement over time and plots it as such.

What that 47.6 percent measures deserves stating precisely. It is the fraction of tasks where an expert, shown two finished artefacts, could not say the human one was better. It is a judgement on the artefact, not on the process. The task was specified up front in full, the model got one shot, and nobody came back with a clarifying question or a change of scope. The authors say this themselves in the limitations. Tasks are precisely specified and one shot, not interactive, and the benchmark covers self-contained computer work rather than anything involving tacit knowledge or proprietary tooling.

That framing makes the result both more and less impressive than it sounds. More, because a blind expert preferring or accepting a model artefact half the time on real professional deliverables is well beyond what anyone would have predicted three years ago. Less, because a large share of professional work is the part before the artefact exists, and that part was done by the human who wrote the task.

Horizontal bar chart of win-or-tie rate against the human deliverable on the GDPval gold subset: GPT-4o 12.5%, o4-mini 29%, o3 35%, GPT-5 39%, Claude Opus 4.1 47.6%, with a dashed line marking 50% parity.
Win-or-tie rate against the human deliverable, GDPval gold subset. Claude Opus 4.1 leads at 47.6%, just short of parity.

How the grading holds up

The paper reports inter-rater agreement between human graders of 71 percent, and agreement between the automated grader and human graders of 66 percent. That the automated grader is within five points of human-to-human agreement is the argument for using it, and it is a reasonable one. It also means about three in ten pairwise calls are contested even between experts, so a few points of difference between two models near the top should not be read as a ranking.

We would have liked to see agreement broken down by occupation. Pairwise preference on a legal memo and pairwise preference on a set of engineering drawings are different kinds of judgement, and a single pooled agreement number hides where the instrument is sharp and where it is blunt.

Where the 100x came from

The multiples that spread after the release are in Table 2 of the paper, in a row the authors label the naive ratio. Expert completion of a task took 404 minutes on average and cost 361 dollars at the occupation's median hourly wage. Divide that by API time and API cost and GPT-5 comes out about 90 times faster and about 474 times cheaper. GPT-4o, which is faster and cheaper per token but wins far less often, shows 327x and 5,172x on the same naive basis. Those numbers are the source of the headline, and the paper's own next sentence is that once you incorporate time to review and redo work, the payoff shrinks.

The realistic scenarios the authors model assume an expert reviews the model output, which took 109 minutes on average, and completes the work by hand when the output is unsatisfactory. Under the scenario of trying once and then fixing, GPT-5 is 1.12 times faster and 1.18 times cheaper than an unaided expert. Under trying several times before fixing, it is 1.39 times faster and 1.63 times cheaper. Those are the honest numbers for what the paper can support, and they are modest.

Even those figures are generous in two ways the authors flag. The analysis does not charge the human baseline for review time, though real professional deliverables get reviewed too. And it does not price catastrophic mistakes, which in some of the 44 occupations are the whole reason the review exists. A one in ten chance of a disastrous error is not captured by a 1.18x saving.

Two bar charts showing GPT-5's speed and cost multiple over an unaided expert. Naive ratio (API time and cost only): 90x faster, 474x cheaper. Realistic, trying once then fixing: 1.12x faster, 1.18x cheaper. Realistic, trying several times then fixing: 1.39x faster, 1.63x cheaper.
The multiple that travelled with the release, and the multiple the paper's own realistic scenarios support, GDPval Table 2.

What we take from it

GDPval is a serious attempt to measure something evaluation has mostly avoided, and the choice to have experts compare finished artefacts blind is the right one for the question it asks. The result we trust is the win-or-tie rate. The result we do not trust is any speed or cost multiple that leaves out the reviewer, because the paper itself shows that including the reviewer collapses two orders of magnitude to a fraction of one.

The experiment we want next is the interactive version. Give the model the same task but let the grader ask for revisions the way a manager would, and count the rounds and the review minutes. If the win-or-tie rate holds up under iteration and the review time drops as reviewers learn what to check, the economics will move toward the naive ratio on their own. If it does not, then the benchmark has measured a real but narrow capability, producing a good artefact from a complete specification, and the 47.6 percent figure should be quoted with that qualifier attached.

Sources

  1. Patwardhan et al., GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks (OpenAI)
  2. GDPval paper on arXiv (2510.04374)