o3, ARC-AGI, and what it means to pass a benchmark at $4,560 a task
o3 scored 87.5 percent on the ARC-AGI-1 semi-private set, above the 85 percent target, using 1,024 samples per task and around $456,000 of compute. The score is real. The question is what unit it should be reported in.
Two scores, not one
ARC Prize published o3's results on Friday and the important thing about the announcement is that there are two sets of numbers. In the high efficiency configuration, six samples per task, o3 scored 75.7 percent on the 100 task semi-private evaluation for a total of about $2,680, or $26 per task. In the low efficiency configuration, 1,024 samples per task, it scored 87.5 percent for roughly $456,000, which works out to $4,560 per task. On the 400 task public set the equivalent figures are 82.8 percent at $167 per task and 91.5 percent at $1,900 per task.
For context, GPT-3 scored 0 on this benchmark in 2020, GPT-4 was near 0, and GPT-4o managed 5 percent earlier this year. The best private set score from the 2024 Kaggle competition, with its compute limits, was 55.5 percent, up from 33 percent at the start of the year according to the ARC Prize technical report. So a jump to 75 percent at $26 a task is a large result on its own terms, before anyone spends a house on the second configuration.
Francois Chollet, who created the benchmark, called it a genuine breakthrough and a qualitative shift, and in the same post said he does not think o3 is AGI and that passing ARC-AGI does not equate to achieving it. He also noted that o3 still fails some very easy tasks. We think all of those statements are compatible, and the reason they are compatible is the cost axis.
Cost as a first class axis of the result
ARC Prize reports every score alongside its cost per task, and the public leaderboard only lists systems that ran for under $10,000 total. The high efficiency o3 run qualifies for the leaderboard under that rule. The 87.5 percent run does not. It is reported, but as an out of budget data point, and that distinction is doing a lot of work.
The reason it matters is that 1,024 samples per task is a knob you can keep turning. If the score at 6 samples is 75.7 and at 1,024 it is 87.5, then the benchmark is measuring a curve, and any single number on that curve is meaningless without its x coordinate. A benchmark result reported without compute is like a top speed reported without saying whether the car was going downhill.
This is not new in principle. Pass@k has always been a curve. What is new is that the gap between the cheap point and the expensive point is now 12 points on a benchmark people care about, and the expensive point costs more than most academic groups spend on compute in a year. When the top of the leaderboard depends on a budget only a handful of organisations have, the leaderboard stops being a comparison of methods and becomes a comparison of bank accounts.
What the 2024 competition learned about the dataset
The ARC Prize technical report is candid about the limits of ARC-AGI-1 itself. The private set score went from 33 to 55.5 percent this year, mostly through deep learning guided program synthesis and test-time training, and the report says explicitly that the dataset has limitations the organisers now understand better. Five years of public exposure means the task distribution has been studied closely, and test-time training on the demonstration pairs of each task turns out to be a strong lever.
That matters for reading the o3 number, because the semi-private set is the one the Kaggle entrants could not see, and 87.5 on it is well above the 55.5 that the best compute-limited entry reached. The gap between those two figures is a mix of method and budget, and the report does not let you separate the two. What it does establish is that the 85 percent target was set in 2024 against a public dataset five years old, and that reaching it was a matter of when rather than whether.
What o3 shows, in Chollet's description, is a form of deep learning guided program search at a scale nobody else can run. That is a real capability. It also means the benchmark has been approached from the direction its designers hoped for, searching over programs rather than pattern matching from a fixed corpus, and the remaining question is whether the search is efficient enough to count as the kind of generalisation the benchmark was meant to measure.
How the organisers responded
The response is ARC-AGI-2, which ARC Prize says is coming with the 2025 competition. Early testing suggested o3 might land under 30 percent on it while humans exceed 95. When the benchmark launched in March 2025 with 120 calibrated tasks per set, the numbers came in lower than that. o3 preview in low compute mode scored about 4 percent at $200 per task, o1 pro about 1 percent, and o3 mini high 0.0 at $0.41 per task. The human panel hit 100 percent with at least two solvers per task, and averaged 60 percent, at a cost the organisers put at $17 per task.
The efficiency rules got stricter too. The 2025 Kaggle competition gives roughly $50 of compute per submission, double the previous year, with no internet access. The $700,000 grand prize unlocks at 85 percent. Every ARC-AGI-2 task was solved by at least two humans in two attempts or fewer in a controlled study. So the benchmark now ships with a human cost baseline as well as a human accuracy baseline, which is the right response to a result that was bought as much as earned.
What we would want next is for every benchmark that reports a frontier score to adopt the cost per task column. Not as a footnote, but on the same axis as accuracy, with the human cost on the same chart. The o3 result is impressive. Reporting it as 87.5 percent and stopping there is the least informative way to describe it.
Sources
From the foundation