Naked accuracy is marketing: the cost axis arrives on leaderboards
ARC Prize ran every major reasoning system through the same tasks and found no single winner, only a frontier stretching from $200 a task down to four cents. Why a score without a price is now an incomplete result.
A test with no winner
On 5 June Mike Knoop at ARC Prize published the results of running thirteen reasoning systems through the ARC-AGI semi-private sets. The list covered o3-preview, o3, o3-pro and o4-mini from OpenAI, Claude Sonnet 4 and Opus 4 from Anthropic, Gemini 2.5 Flash and Pro from Google, DeepSeek R1, Grok 3 and Grok 3 Mini, Llama 4 Maverick, and GPT-4.1 Nano. The headline in the post is that there is no headline. No system won. What came out instead was a Pareto frontier on accuracy against cost, and every lab has a point somewhere on it.
The spread on ARC-AGI-1 is the part we keep coming back to. o3-preview at low compute scored 75.7 percent at about $200 per task. o3 at high compute scored 60.8 percent at $0.50 per task. Gemini 2.5 Flash scored 33.3 percent at $0.037 per task. Between the top and bottom of that table the accuracy roughly halves and the price drops by a factor of about five thousand. Any ranking that reports the first column and hides the second is telling you something closer to an advertisement than a measurement.
What the frontier actually looks like
Read down the ARC-AGI-1 results and the shape of the trade becomes clear. o3-pro at high compute scored 59.3 percent for $4.16 a task, which is worse than plain o3 on both axes. o4-mini at high compute hit 58.7 percent for $0.406, nearly the same accuracy as o3 for a bit less money. Claude Sonnet 4 reached 40 percent for $0.366 and Claude Opus 4 reached 35.7 percent for $1.25, so within one lab the cheaper model was also the more accurate one on this benchmark. Gemini 2.5 Pro landed at 33.0 percent for $0.569, almost tied with Flash on accuracy at fifteen times the cost.
Those last two comparisons are the useful ones. If you only had accuracy, Opus 4 and Sonnet 4 look like neighbours and Gemini Pro and Flash look like twins. With cost in the picture, one member of each pair is simply dominated on this task. That is information a buyer needs and a single number cannot carry.
ARC-AGI-2 is harder and flatter. Claude Opus 4 led at 8.6 percent for $1.93 a task, o3 at high compute scored 6.5 percent for $0.834, o4-mini scored 6.1 percent for $0.856, and Claude Sonnet 4 scored 5.9 percent for $0.486. Nobody is close to solving it, and the ARC Prize post says as much in plain words, that new ideas are still needed. But even down at single digits the cost column changes the story. Opus 4 is the most accurate and also four times the price of Sonnet 4 for under three extra points.
Why cost stopped being optional
For most of the history of benchmarks the cost of a run was roughly the same for every model of a given size, so leaving it out lost little. Reasoning systems broke that assumption. ARC Prize identifies three test-time techniques the labs are now using: long-running inference where the model keeps generating tokens for as long as it likes, recomposition of knowledge by remixing chains of thought, and parallel sampling with some form of voting. All three convert money into accuracy at inference time. A model with a thinking budget knob does not have one score. It has a curve.
The o3-preview number is the cleanest illustration. That 75.7 percent was bought for roughly $200 per task, and the ARC Prize team ran it under their own testing rather than trusting a vendor claim. The same model family, tuned for release, delivered 60.8 percent for fifty cents. Both numbers are true. Reporting either one on its own, without the price, invites the reader to draw the wrong conclusion about what a paying customer will get.
This also explains why we have stopped trusting vendor benchmark tables that omit test-time settings. If a lab can move a score by fifteen points by changing a compute setting and a budget, the setting and the budget are part of the result. A table that hides them is not wrong, exactly. It is incomplete in a way that happens to favour whoever wrote it.
What we are changing in our own reporting
We ran into this ourselves earlier in the year while comparing open reasoning models on internal tasks. Our first table had accuracy only and looked decisive. Once we added dollars per task at published API prices, two of our conclusions reversed and one model we had written off turned out to be the sensible default for anything run at volume. The exercise cost an afternoon and changed the recommendation we sent to a partner.
From this month, any evaluation we publish that involves a model with a configurable reasoning budget will report cost per task alongside accuracy, using the ARC Prize convention of a semi-private set and the prices in force on the day of the run. Where we cannot get a price, because the model is local or the vendor does not publish one, we will report tokens generated per task instead, which is the quantity the price is built from anyway. A bare accuracy number will not appear in our tables without one of those two companions.
We expect the field to move the same way, and quickly. Once one respected evaluation publishes a two-axis chart, a single-axis chart from anyone else starts to look like a choice to hide something. The interesting question is what the frontier looks like in a year. Our guess is that the top-left corner moves up faster than the bottom-right corner moves left, and that the four-cent tier will be where most real work gets done. We would like to see someone test that guess with the same rigour ARC Prize just applied.
Sources
From the foundation