Two numbers, five weeks apart

In September OpenAI published GDPval and reported that Claude Opus 4.1 produced a deliverable that expert graders rated as good as or better than a human professional on 47.6 percent of a 220 task gold set. GPT-5 came in at 39.0 percent and o3 at 35.2 percent. The paper described models as beginning to approach parity with industry experts. On 30 October a group of 47 authors led by Mantas Mazeika, with the Center for AI Safety and Scale AI, posted the Remote Labor Index. Its best performing agent, Manus, automated 2.5 percent of projects. Grok 4 and Claude Sonnet 4.5 reached 2.1 percent, GPT-5 1.7 percent, ChatGPT Agent 1.3 percent, and Gemini 2.5 Pro 0.8 percent.

We have spent the past week with both papers, and the disagreement is smaller than the headlines suggest. The two benchmarks measure different things, grade them differently, and hand the model different amounts of help. Once you line those up, an order of magnitude gap is roughly what you would expect.

What each benchmark actually asks

GDPval is built from 1,320 tasks across 44 occupations in the nine sectors that contribute most to United States GDP. Each task was written by an industry professional with an average of 14 years of experience, and the expert reference deliverable took seven to nine and a half hours to produce. The open sourced gold subset has five tasks per occupation. A model receives the task prompt and any reference files, produces a deliverable, and that deliverable is compared blind against the human one by another expert in a pairwise judgement that averaged over an hour.

The Remote Labor Index is built from 240 projects in 23 categories of remote work, contributed mostly by 358 verified Upwork professionals from work they had already completed and been paid for. The projects total more than 6,000 hours of labour worth roughly 140,000 dollars, with an average price of 632.60 dollars and a median of 200 dollars. Average completion time by the original freelancer was 28.9 hours, median 11.5. The deliverables include CAD files, video, audio, and 3D models, and the authors estimate the projects are more than twice as complex as those in prior benchmarks.

So the unit of work differs by a factor of three or four in hours, and the RLI projects come with the full messiness of a client brief rather than a task written for the purpose of evaluation.

Win or tie is a different metric from automated

The larger difference is in the scoring. GDPval reports the rate at which a grader preferred the model output or called it a tie. A tie counts as a success, and the comparison is relative to one specific human deliverable. A model output that is mediocre but comparable to a mediocre human output scores a win or tie.

RLI asks a different question. Three evaluators independently judge whether the deliverable would be accepted by a reasonable client as commissioned work, write a justification, and the majority decides. Inter-annotator agreement was 94.4 percent. That is an absolute standard, and it is the standard a freelancer is actually held to. The failure breakdown shows what trips agents up. Poor quality was cited in 45.6 percent of failed deliverables, incomplete work in 35.7 percent, corrupted files in 17.6 percent, and internal inconsistencies in 14.8 percent.

GDPval's own failure analysis points the same direction. The most common reason experts preferred the human deliverable was instruction following, with formatting errors and incomplete deliverables next. Those are exactly the categories that push a project from "comparable to a human draft" to "a client would send it back". The two papers agree on where models fail. They differ on whether a near miss counts.

The agent harness matters as much as the model

GDPval evaluates a model producing a deliverable from a prompt and reference files. RLI evaluates an agent that has to find its way through a multi hour project end to end, handle file formats, and hand back something a client can open. The 17.6 percent of failures attributed to corrupted files are a harness problem more than a reasoning problem, and they would not show up in a benchmark where output is a document graded by a person.

The RLI Elo table makes the same point from another angle. With the human baseline anchored at 1,000, Manus scored 509.9, Grok 4 468.2, ChatGPT Agent 454.3, Sonnet 4.5 441.7 and GPT-5 436.7. ChatGPT Agent ranks above Sonnet 4.5 on Elo while scoring lower on automation rate. Relative quality and absolute acceptability are separate axes, and an agent can move up one without moving up the other.

Which number to use for which question

If the question is whether a frontier model can produce work of professional quality on a well specified task when a human is going to review it, GDPval is the better instrument, and its answer is that the best models are competitive with an expert draft somewhere around half the time. The paper's own framing of "try the model first, fix if needed" is the honest reading. That is a story about assistance.

If the question is whether an agent can replace a freelancer on a real project without anyone checking, RLI is the right instrument, and the answer in October 2025 is between one and three percent. That is the story about automation, and it is the one economists and policy people actually want.

What we would want next is the same projects graded both ways. Take the RLI set, produce model deliverables, and have experts do GDPval style pairwise comparisons against the original freelancer's work alongside the accept or reject judgement. Our guess is that the pairwise win or tie rate on RLI would land far below 47.6 percent because of project length, but well above 2.5 percent. That gap, measured on one dataset, would tell us how much of the distance between the two headline numbers is the metric and how much is the work.

Sources

  1. Mazeika et al., Remote Labor Index: Measuring AI Automation of Remote Work (arXiv 2510.26787)
  2. Remote Labor Index, full text
  3. GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks (arXiv 2510.04374)
  4. Scale AI research papers