ARC Prize launches: a million dollars for a benchmark AI could not pass
Francois Chollet and Mike Knoop have put a prize pool of over a million dollars on ARC-AGI, where the state of the art sits at 34 percent. A look at the prize as an evaluation design, and at the assumptions it will test.
What was announced
On June 11 Mike Knoop and Francois Chollet announced ARC Prize, a competition with a prize pool of more than a million dollars for anyone who can beat and open-source a solution to the ARC-AGI evaluation. Chollet introduced the benchmark in 2019 in his paper On the Measure of Intelligence. At the time the best score was 20 percent. Five years and several generations of language models later, the announcement puts the state of the art at 34 percent.
The framing is combative on purpose. The launch post says that large language models are great memorisation engines, that they memorise reasoning patterns and apply them in adjacent contexts, and that they cannot generate new reasoning for novel situations. ARC-AGI was built to resist memorisation, and the prize exists to find out whether anything can pass it.
What the benchmark is
Each ARC task is a handful of input and output grids, usually around three demonstration pairs, followed by a test input. The solver has to infer the rule from the demonstrations and produce the test output. There is no instruction text. The rules draw on things like symmetry, counting, object persistence and simple physics, which Chollet argues are the core knowledge priors any human has and no training set can fully enumerate.
The public training set has 400 tasks and the public evaluation set has another 400. The competition is scored on a private set of 100 tasks that has never been published, and the organisers have added a semi-private set of 100 for testing commercial APIs that they cannot run offline. The private set is the whole design. A benchmark that lives on the internet ends up in the next crawl, and a memorisation engine will then score well on it for reasons that have nothing to do with reasoning.
The prize as an experiment
It is more useful to read the prize as an evaluation experiment than as a contest. The organisers have a hypothesis, that current systems memorise rather than reason, and they have designed conditions under which the hypothesis can be refuted. The conditions are what make it interesting.
First, the Kaggle track runs offline. Solutions get a fixed compute budget on the competition hardware, with no internet access, so a frontier API cannot be called at test time. Second, prize-winning solutions must be open sourced, so a winning score can be reproduced rather than taken on trust. Third, the target is set far above the current state of the art. The organisers describe the tasks as easy for humans and impossible for modern AI, and they are putting real money on the word impossible.
Those three constraints decide what kind of result would count. A closed model hitting a high score on the public set would prove nothing under these rules, because the public set is contaminated by construction and the model's training data is unknown. A reproducible offline solution on the private set is a different kind of evidence, and it is the only kind the prize accepts.
Where we think the design is vulnerable
The compute cap is the part we would watch. It is there to stop people brute-forcing tasks with money, and that is a sensible thing to stop. But it also means the competition cannot distinguish between 'this task is beyond current methods' and 'this task is beyond current methods at this budget'. If someone demonstrates a high score with a hundred times the compute, the prize's rules will call that a non-result while the wider field will call it progress. The organisers will then have to decide which claim they were making.
The second vulnerability is the memorisation argument itself. It is stated as a property of language models, but the benchmark can only measure behaviour. If a system fine-tuned on ARC-style tasks scores well, the organisers can say it memorised the distribution of tasks, and the entrant can say it learned the priors. The private set makes that argument harder to have dishonestly, and it does not make it go away.
The third is the human baseline. The launch material says the tasks are easy for humans, and the tasks we have tried are, but a benchmark defined by the gap between humans and machines needs a measured human number, not a description. A prize this size should publish one.
What would change our mind
We go into this expecting the grand prize to go unclaimed this year and the top offline score to move by ten or fifteen points through search and program synthesis methods that owe little to large language models. If instead a language model based approach lands near the top of the private leaderboard under the compute cap, that is evidence against the memorisation story, and we would take it seriously.
The more interesting outcome is the one the rules do not cover. If a closed frontier system posts a very high score on the semi-private set with unbounded compute, the prize will have shown that its target was reachable and that its constraints were what kept it out of reach. That would be a result about evaluation design as much as about intelligence, and it is the one we would most want to see argued over.
Sources
From the foundation