ARC Prize Verified: the benchmark that hired auditors
ARC Prize now runs frontier systems on its hidden test set, has four academics sign off on the method, and takes money from labs on the condition that the money cannot touch their scores. Notes on self-reported benchmarks as a market failure and whether this fix can scale.
What was announced
On November 4 the ARC Prize Foundation announced a program called ARC Prize Verified. Under it, a frontier system gets a score on the official leaderboard only if the foundation itself ran the system against the hidden test set, and only if the methodology has been documented and validated by an academic panel. Verified scores carry a badge. Scores that come from a lab running the public tasks on its own hardware do not.
The panel is four people: Todd Gureckis at NYU, Guy Van den Broeck at UCLA, Melanie Mitchell at the Santa Fe Institute, and Vishal Misra at Columbia. The money comes from five labs, Ndea, xAI, Google, Nous and Prime Intellect, on top of earlier individual donors including Tyler Cowen, Dharmesh Shah and Aaron Levie. The announcement is explicit that donating does not affect how a donor is scored. That one sentence is doing more work than anything else in the post.
Why self-reported scores were failing
Since ARC-AGI-2 launched in March, the foundation has published a cost per task alongside every score, and the two axes together are what make the leaderboard useful. o3-preview at low effort was estimated at 4 percent for around 200 dollars a task. The 2024 Kaggle winner sat at 3 percent for 25 cents. A human panel averaged 60 percent at 17 dollars a task. Those numbers only mean something if the same people ran all of them on the same tasks, and the hidden set is the only place that condition holds.
The problem with a self-reported score is a plain incentive one. The lab chooses the tasks, the scaffold, the number of samples, the effort setting and the moment to stop. None of that is fraud. It is just that each choice is made by the party with the most to gain from a high number, and the reader cannot see any of the choices. Model cards from OpenAI, xAI, Anthropic and Google DeepMind now cite ARC, so the number has become part of a product launch rather than a research result. When the number sells the product, nobody on the selling side has a reason to make it smaller.
Economists would call this a market for lemons. Buyers of a benchmark score cannot tell a careful measurement from an optimistic one, so they discount all of them, and the careful lab gets no credit for its care. The standard repair is a trusted third party who can inspect what buyers cannot. That is an auditor. ARC Prize has now hired four.
What the funding structure gets right
The interesting design choice is who pays. The labs whose scores get audited are the ones funding the audit. That looks like a conflict until you notice that the alternative, a benchmark funded by nobody, ends in the benchmark going away or in scores that cost the foundation 200 dollars a task to produce being run only when someone else covers the bill. Auditing frontier systems on hundreds of hidden tasks is not cheap, and the labs are the only parties with both the money and the interest.
The announcement says the new lab funds go to three things: more ARC-AGI-3 games at higher quality, technical infrastructure such as an API and an offline engine, and wider researcher access. None of that is verification itself, which is consistent with the promise that donor money does not touch scoring. We would like to see that promise written down in more detail, ideally with the hidden set custody and the panel's sign-off procedure described somewhere a sceptic can read. A one-line assurance is a start and not an audit trail.
Can this scale past one benchmark
ARC is unusually well suited to this arrangement. The tasks are small, the answer is a grid, grading is exact, and the foundation already owned a private set that nobody had seen. Most benchmarks that matter today have none of those properties. Coding benchmarks need a container per task and a judgment call about what counts as passing. Agentic benchmarks need live environments that drift. Long-form evaluations need graders, and graders need calibration. The cost of independent verification rises steeply with each of those.
There is also a supply problem. Four professors can validate one foundation's methodology for one benchmark. The field runs on dozens of benchmarks and the labs release models every few weeks. If verification is going to be the norm rather than a badge on one leaderboard, it needs to be a service with staff and a fee schedule, closer to a testing lab than a volunteer panel. We do not know who builds that, and we doubt the labs will fund a general one as readily as they funded this one, because a general one would verify the scores they least want verified.
What we would want someone to try next is a second benchmark adopting the same rules, with a different panel and a different set of funders, so we can see whether the ARC arrangement was a property of ARC or a template. If the second one works, the sentence 'We evaluated on X' starts to carry a footnote about who ran it, and that footnote would be the most useful thing to happen to evaluation in years.
Sources
From the foundation