458 people in a room: how ARC-AGI-3 measured its human baseline
ARC Prize paid members of the public to play its interactive environments once each, first try, and then anchored the AI score to the median human. Notes on why the human side of a benchmark is usually the sloppier measurement, and what a properly collected one changes.
The half of the benchmark nobody audits
Every benchmark that reports a human baseline has two measurements in it, and the human one almost never gets the scrutiny the model one does. Model runs are logged, seeded, re-run and argued over on forums. The human number is often a handful of authors or contractors solving the tasks at their desk, sometimes after having written them, and it gets quoted for years as if it were a physical constant.
ARC Prize published the human side of ARC-AGI-3 on April 14, and it is the first time we have seen a benchmark team treat the human measurement with the same care as the model measurement. The headline figures are 458 participants, 135 environments, and a protocol where every person saw each environment exactly once and got exactly one attempt. This post is about why those design choices matter more than the numbers they produced.
How the data was collected
The participants were members of the general population recruited through weekly in-person focus groups at a testing centre in San Francisco. ARC Prize describes the group as spanning various levels of education, income, job sector and age, which is a different population from the authors of the benchmark or from the Kaggle competitors who will try to beat it. Compensation was a base payment of about 130 dollars plus 5 dollars per environment solved, so there was an incentive to try hard without an incentive to game the protocol.
The first-run rule is the piece that makes the comparison honest. ARC-AGI-3 environments are interactive games with hidden rules that the player has to discover by acting. If a person plays the same game twice, the second run measures memory of the solution, not the ability to find it. An AI agent evaluated on the benchmark sees each environment cold. By giving humans a single cold attempt too, both sides of the comparison receive identical information.
What was recorded is the full step-by-step interaction, not just pass or fail. The public release covers 342 human sessions across 25 public environments, with per-environment completion rates, action counts and efficiency distributions. The remaining environments are held back for the competition, which is the right call, but it means the public replay data is a sample rather than the whole set.
Why the median, and why the cap moved
The scoring anchor is the interesting decision. An AI's score on a level is its efficiency relative to human efficiency, where efficiency is roughly how few actions it took to win. The earlier scoring used the second-best human as the reference point. The new scoring uses the median human per level.
The reasoning is variance. The second-best player on a level is partly a measure of skill and partly a measure of luck, since among a few dozen people someone will stumble onto the rule early by chance. Anchoring to that person makes the benchmark's difficulty a function of who happened to show up on a given Tuesday. The median is stable under adding or removing a few outliers, and it describes the typical person rather than the best one.
With the median as anchor, an AI can beat a human on a level, so the per-level score cap rose from 100 percent to 115 percent. Without a cap above 100, an agent that did well on most levels and badly on one would be dragged down by the single weak result with no way to compensate. Fifteen percent of headroom is a judgement call, and we would like to see the sensitivity analysis behind it, but the principle is sound.
What solvable means
ARC Prize states that every one of the 135 environments was beaten by at least two humans and typically by five or more, and on that basis calls the benchmark 100 percent solvable by humans. We want to point out how much weaker and more honest that claim is than the usual version. It does not say the median person solves every level. It says that for every level, there exist ordinary people who solve it on their first try. The per-level completion rates in the released data will show how far below 100 percent the typical person sits, and those rates are the number a model score should be read against.
This is the part we think other benchmark authors should copy. A task that only its designers can solve is not a test of general ability, and a task that a hundred random people fail is telling you something about the task. Publishing the completion distribution rather than a single pass rate lets a reader tell the two apart.
What a good baseline changes
A carefully collected human baseline changes what a model score means. When the human number is soft, a model at 30 percent of human might be at 60 percent of a properly measured median or 15 percent of it, and nobody can say which. When the baseline is 458 first-run sessions from a mixed population, the model number is a real ratio and it can be tracked across model releases without the denominator shifting under it.
It also changes what counts as a fair fight. The protocol here removed practice, hints and repeated attempts from the human side so that neither party had an advantage in information. If a lab then evaluates an agent with retries or with environment descriptions in the prompt, the comparison breaks, and now the breakage is visible because the human protocol was written down.
What we would want next is the same protocol run in a second city with a second population. ARC Prize calls the human data the receipt for the capability gap. A receipt is more convincing when two shops print the same total.
Sources
From the foundation