Where it started

ARC-AGI-2 launched on 24 March 2025, after the original ARC had been pushed past 90 percent by o3 at very high cost per task. The new set had 1,000 training tasks and three evaluation splits of 120 tasks each, public, semi private and private. Every evaluation task had been solved by at least two humans in under two attempts. The human panel score with that rule was 100 percent, and the average panel member scored 60 percent, at a cost the organisers estimated at 17 dollars per task.

Frontier models at launch were near zero. o3-preview at low compute scored 4 percent at about 200 dollars per task. o1-pro scored 1 percent at the same price. o3-mini-high and GPT-4.5 scored 0.0 percent. The ARC Prize 2025 grand prize of 700,000 dollars required 85 percent under Kaggle compute limits. The design paper, published in May, described the tasks as demanding multi rule compositional reasoning, multi step sequential reasoning, contextual rule application, and in context definition of symbols. It reported 407 human participants across 515 sessions, with 75 percent of task attempts succeeding and a median solve time of just over two minutes.

The year in between

By the May paper the semi private leaderboard had o3 at medium compute and o3-mini-high both at 3.0 percent, o4-mini at 2.4 percent, and Claude 3.7 at 0.9 percent. The Kaggle competition, with its 50 dollar budget for the 120 private tasks, closed in November with the top team, NVARC, at 24.03 percent using synthetic data generation. Nobody came near 85 percent under the compute cap.

The frontier lab numbers moved faster. When ARC Prize published its 2025 results in December, Claude Opus 4.5 with 64k thinking tokens had a verified 37.6 percent at 2.20 dollars per task. Gemini 3 Pro alone scored 31 percent at 81 cents per task, and wrapped in the Poetiq refinement harness it reached 54 percent at about 30 dollars per task. All four major labs, OpenAI, Anthropic, Google DeepMind and xAI, were reporting on ARC-AGI-2 by then. That was the state of play at the end of 2025. Nine months from 4 percent to 54 percent, though the 54 percent needed a scaffold and a hundred times the per task cost of the base model.

February 2026

This month Google released Gemini 3.1 Pro and its model page lists a 77.1 percent score on ARC-AGI-2, in the ARC Prize Verified category. The same table gives Gemini 3 Pro at 31.1 percent, Claude Sonnet at 58.3 percent and Claude Opus at 68.8 percent. Other reporting puts GPT-5.2 at 52.9 percent. We have not found a verified cost per task figure for the 3.1 Pro run and we are not going to guess one.

So a benchmark that opened with the best model at 4 percent has the best model at 77 percent eleven months later, and a second model above two thirds. That is not 85 percent, and it is not under the Kaggle compute limit, so the grand prize condition has not been met. But the design goal stated in the launch post, a test that stays easy for humans and hard for AI, has been overtaken. The average human panel member scored 60 percent. The best model now scores higher than the average human.

What the short life tells us

ARC-AGI-2 was built by people who had watched ARC-AGI-1 fall and who tried to close the specific loopholes that killed it. It fell anyway, and it fell faster. We take three things from that.

The first is that "easy for humans, hard for AI" is a property of a moment, not of a task. The set was chosen by filtering out what 2024 models could do. That is a description of the gap at filtering time. The tasks did not change. The models did.

The second is that the numbers most worth tracking were the cost figures. At launch, human quality cost 17 dollars per task and machine quality at 4 percent cost 200. By December a 37.6 percent solution cost 2.20 dollars. The capability curve and the cost curve moved together, and the cost curve is the one that tells you whether a benchmark score means anything outside a leaderboard.

The third is that ARC Prize seems to have drawn the same conclusion. The December post announced ARC-AGI-3 for early 2026 with an interactive format that tests exploration, planning, memory and goal acquisition rather than static puzzles. Whether an interactive test resists the same trajectory is the open question, and given the history we would not bet on more than a couple of years.

What we would want from the next one

If we were designing the successor we would publish two things alongside every task. The first is the date the task was written and the model generation it was filtered against, so that a future reader can tell how much of the difficulty was contemporaneous. The second is a standing cost reporting requirement, so that a score without a dollars per task figure is not accepted as verified. ARC Prize already does a version of the second. Making it mandatory would keep the benchmark honest for whatever life it has.

Sources

  1. ARC Prize: Announcing ARC-AGI-2 and ARC Prize 2025
  2. Chollet et al., ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems (arXiv 2505.11831)
  3. ARC-AGI-2 paper, full text
  4. ARC Prize 2025 Results and Analysis
  5. Google DeepMind: Gemini 3.1 Pro model page (benchmark table)
  6. MarkTechPost: Top 7 benchmarks for agentic reasoning (April 2026)