When the score stops carrying information

ARC-AGI-3 is a set of 135 novel interactive environments, each one a small game with unstated rules, and each one solved by at least two untrained humans before it was admitted to the benchmark. Against the semi-private set, GPT-5.5 scored 0.43 percent and Opus 4.7 scored 0.18 percent. More than a million games have now been played on the benchmark, so those are not small-sample numbers. They are simply very small numbers.

A difference of a quarter of a point between two models at that floor tells you almost nothing about which is better at the thing the benchmark is meant to test. It could be one level on one game. The ARC Prize team's response was to stop looking at the aggregate and download the logs. They pulled reasoning traces and action histories from public runs, established the intended strategy for each game by hand, used automated tools to flag candidate failure patterns, and then cross-validated by human review across both models. The published analysis covers 160 replays.

Three ways to fail a game you have never seen

The first pattern they call local observation with global incomprehension. A model notices that a specific action has a specific effect, ACTION3 rotates the object, and records it correctly, then never assembles those observations into a model of the game that could guide a plan. In one environment, cd82, Opus identified the rotation correctly and then never used it in its strategy for orienting the bucket that the level needed. The observation was right and it went nowhere.

The second is misaligned abstraction. Faced with a grid of coloured cells and a few controls, the models reach for a game they already know. Tetris, Frogger, Sokoban, Breakout. The analogy then decides what the model tries. In ls20, GPT-5.5 repeatedly classified the game as Breakout and kept testing Breakout affordances, which meant it never discovered the key combination the level actually required. Nothing in the environment said Breakout. The model brought the label with it and the label closed off the search.

The third they call false victories. A model completes level one with a wrong theory of the mechanics, because the wrong theory happens to produce the right actions on that level, and the success hardens the theory. When level two requires the real mechanic, the model cannot let go. In ka59, Opus won the first level on a mistaken theory of teleportation and then failed the second. On a scoreboard that first level counts as progress. In the replay it is the cause of the failure that follows.

Two models, two different mistakes

The most useful finding for anyone building on these models is that they fail differently. Opus was better at short-horizon mechanic discovery, at working out what a button does in a few moves, and then latched onto false invariants aggressively. GPT-5.5 generated a wider range of hypotheses and had more trouble converting an observation into a committed plan. The team's summary is wrong compression against failure to compress. One model builds a world model too early and cannot revise it. The other keeps its options open and never acts on them.

That distinction is invisible in the scores, where the two models are separated by a quarter of a percentage point. It is exactly the kind of thing a developer choosing between the models for an agentic task would want to know, and it came from reading traces rather than from counting wins.

What process-level evaluation costs

The obvious objection is that 160 hand-reviewed replays do not scale, and that a benchmark whose interesting output requires a research team to read logs is a benchmark that only its authors can interpret. That is half right. Hand review established the ground truth and the taxonomy. The automated flagging did the search. Once the three patterns are named and there are labelled examples of each, a classifier over traces can look for them in the next million games without a human in the loop, and the human effort becomes a validation step rather than the method.

The deeper point is that the score and the replay answer different questions. The score says whether the model reached the goal. The replay says whether it understood the mechanic, and a level can be won without understanding, as ka59 shows, or lost after a correct observation, as cd82 shows. When the benchmark is designed to test the ability to learn a new rule from interaction, understanding is the quantity of interest and the score is a noisy proxy for it. At 0.43 percent the proxy has run out of resolution. At 43 percent it would still be hiding false victories.

What we would like to see

We would like ARC Prize to publish the trace-level labels alongside the scores, so that a model's report card on ARC-AGI-3 reads as a distribution over failure modes rather than a single number. False victories in particular ought to be counted and subtracted, because a level won on a wrong theory is not evidence of the thing the benchmark is measuring.

And we would like someone to test whether the misaligned-abstraction failure can be induced or suppressed. If telling the model in the system prompt that the game is not any game it knows reduces the Breakout errors, that tells us the analogy is a prior it can set aside. If it does not, the prior is deeper than instruction, and that is a more interesting result than any score on the leaderboard this year.

Sources

  1. ARC Prize: ARC-AGI-3 analysis of GPT-5.5 and Opus 4.7 (May 1, 2026)