Nine months, two tables

The launch table for ARC-AGI-2 in March was bleak. o3-preview at low effort was estimated at 4 percent for around 200 dollars a task, o1-pro at 1 percent, and the 2024 Kaggle winner at 3 percent for 25 cents. GPT-4.5 and o3-mini-high scored zero. A human panel averaged 60 percent at 17 dollars a task, and every task had been solved by at least two people. The benchmark was designed to be easy for humans and hard for AI, and in March it was.

The results post published after the competition closed on December 5 reads differently. On the verified semi-private set, Claude Opus 4.5 with 64k thinking tokens scores 37.6 percent at 2.20 dollars a task. Gemini 3 Pro wrapped in a refinement loop built by a company called Poetiq scores 54 percent at about 30 dollars a task. Nine months took the frontier from 4 percent to 54, and the human panel's average is now within reach on cost as well as score.

Where the gain came from

The striking thing is that the largest number on the leaderboard belongs to an application-layer wrapper around a commercial model rather than to a new model. Poetiq's contribution is a loop in which the model proposes a solution, checks it, and revises, and the loop is paid for in tokens, which is why the per-task cost is more than ten times that of the Opus 4.5 entry above it. The results post makes the philosophical version of the point in one sentence: from an information theory perspective, refinement is intelligence.

The Kaggle track, where compute is capped, tells the same story from below. The winning team, NVARC, scored 24.03 percent at 20 cents a task, and the ARChitects, who won in 2024, came second at 16.53. The scores are lower because the budget is lower, but the mechanism is the same. What the model knew before it saw the task matters less than what the loop does with the task once it has it.

The paper prize went to seven million parameters

The 50,000 dollar first paper prize went to Alexia Jolicoeur-Martineau for Less is More: Recursive Reasoning with Tiny Networks, which describes a 7 million parameter model that reaches 45 percent on ARC-AGI-1 and 8 percent on ARC-AGI-2. Third place went to Isaac Liao and Albert Gu for ARC-AGI Without Pretraining, which describes CompressARC, a 76,000 parameter model that scores 20 percent on ARC-AGI-1 with about twenty minutes of compute per puzzle on a consumer RTX 4070. Neither has read the internet. Both do their work at test time.

Put next to the Poetiq result, this is a strange pair of facts to hold at once. A frontier model with a loop around it and a network of a few million parameters, also with a loop around it, are both making progress on the same benchmark by the same route. The base model contributes priors, and the priors help, but the score is produced by iteration. The results post also notes that on one task Gemini 3 Pro used 96 reasoning tokens where Deep Think used 138,000, which suggests even the amount of thinking is a property of the scaffold rather than the weights.

What a score is a score of

This is the part that should worry anyone who reads a model card. ARC-AGI-2 now appears on model cards from OpenAI, xAI, Anthropic and Google DeepMind, and the number on the card is the bare model's number under whatever scaffold the lab chose. The verified leaderboard shows that the scaffold, and the budget it is allowed, can be the difference between the 37.6 percent row and the 54 percent row. So a model card score is a lower bound on what the model can do with a good loop, and an upper bound on what it does without one, and the card rarely says which loop it used.

The cost axis is the only thing keeping this honest. ARC Prize has insisted on a dollars-per-task figure next to every score since March, and it is the cost column that lets you see that the 54 percent system and the 37.6 percent system are different kinds of object. Without it the leaderboard would read as if Gemini 3 Pro had simply beaten Opus 4.5. With it you can see that one model plus an expensive loop beat another model plus a cheap one, which is a claim about scaffolds, and a much more useful one.

What we would want next is for the verified leaderboard to require two entries per system, the bare model at a fixed effort setting and the best scaffold anyone can build around it, each with its cost. The gap between those two rows is the number we actually care about, because it measures how much of the intelligence lives in the weights and how much in the loop. ARC-AGI-3 launches in early 2026 with interactive tasks that need exploration, planning and memory. Our guess is that the loop will matter even more there, and that the field will spend the year arguing about what to call the thing that gets scored.

Sources

  1. ARC Prize, ARC Prize 2025 results and analysis
  2. ARC Prize, Announcing ARC-AGI-2 and ARC Prize 2025 (24 March 2025)