METR's number

METR's Time Horizon 1.1 update, published on 29 January, expands the task suite from 170 to 228 tasks and more than doubles the number of long tasks, those over eight hours, from 14 to 31. The headline metric is the same as before. For each model, find the length of task, measured in how long it takes a human expert, at which the model succeeds half the time. Claude Opus 4.5 comes out at 320 minutes, with a confidence interval of 170 to 729. GPT-5 is at 214 minutes, o3 at 121, Claude Opus 4 at 101.

The trend is the part that gets quoted. Over the full dataset the doubling time is about 196 days, roughly seven months. Restricted to models since 2023 it is 130.8 days with an interval of 107 to 161, and restricted to 2024 onward it is 88.6 days. So a five hour horizon today, on this suite, implies something over a working day within a year if the recent rate holds.

METR is careful about what this rests on, and the caveats are not decorative. The confidence intervals are wide. Only 5 of the 31 long tasks have human baselines, so the top of the curve is estimated more than measured. The task suite was revised substantially, 73 tasks added, 15 removed and 53 updated, and infrastructure changes account for a small but nonzero share of the movement in estimates. None of that undermines the trend. It does mean the point estimates at the long end are soft.

ARC Prize's number

ARC-AGI-3 launched on 25 March. It replaces static puzzles with interactive environments, 135 handcrafted turn based games in total, 25 of them public, the rest semi-private or private. There are no instructions, no stated rules and no stated goal. The agent has to work out what the game wants by acting in it. Scoring is efficiency relative to humans, with each level scored as the square of the ratio of human actions to agent actions, capped at one, and a hard zero if the agent uses more than five times the human action count. The human baseline is the second best of ten testers per level.

On the semi-private set the frontier models scored 0.51 percent in aggregate. Grok 4.20 scored 0, having exceeded the action cutoff on every level. Humans solved 100 percent of environments. The best result from the preview competition, an agent called StochasticGoose that reached 12.58 percent using reinforcement learning with convolutional networks, fell to 0.25 percent when run on the full official benchmark. That last figure is the one we would pay most attention to, because it says the preview environments were learnable and the held out ones were not.

ARC Prize frames the benchmark as measuring exploration, goal acquisition, world model building and continual learning, and as the successor to a generation of benchmarks that reasoning models and coding agents have largely worked through. Two million dollars in prizes are attached across the ARC-AGI-2 and ARC-AGI-3 competitions.

Why both can be true

It is tempting to read these as a contradiction, exponential progress on one side and near zero on the other, and to pick the one that matches your prior. We think that is a mistake, because the two benchmarks share almost no design assumptions and are not measuring the same quantity.

METR's tasks are drawn from software engineering and research work of the kind that is heavily represented in training data and in the post-training that labs do on agents. Success is defined by a checkable output. The task has a specification. The measurement is how long a task, in human terms, a model can carry to a correct result, and it is deliberately designed to track the economically relevant thing, which is whether the model can do work that a person would otherwise be paid hours to do.

ARC-AGI-3 is designed to have none of those properties. There is no specification, no prior distribution the agent can lean on, and the score is penalised quadratically for using more actions than a human. A model that solves a game by trying everything is scored at zero. The benchmark is a measurement of how efficiently a system learns something new from interaction alone, and it is built by people who have spent years removing every foothold that prior knowledge provides.

So the honest reading is that frontier models can now carry long, well specified, in-distribution tasks to completion at rates that double every few months, and cannot yet learn an unfamiliar interactive environment at anything close to human sample efficiency. Both are facts. Neither implies the other.

Which one generalises

The question that matters for anyone planning on these numbers is which measurement predicts performance on the tasks they care about. Our guess is that METR's horizon generalises well to work that looks like METR's tasks and poorly beyond them, and that ARC-AGI-3 generalises well to novel interactive settings and poorly to routine engineering. That sounds like a dodge, and it is really a statement that benchmarks are local.

There is one thing the two agree on, which is the preview to full drop in ARC-AGI-3. An agent that reached 12.58 on the public environments and 0.25 on the held out ones is the same phenomenon that METR sees when it revises its suite and the point estimates move. Performance on what the system has seen overstates performance on what it has not, and the size of the gap is the number you want. METR reports it as wide intervals. ARC Prize reports it as a fifty fold drop.

What we would want next is a cross evaluation. Take the agents that do best on METR's long tasks and run them, with the same scaffolding, on ARC-AGI-3's public set. Take the agents that do best on ARC-AGI-3 and put them on METR's suite. If the rankings agree, the two benchmarks are measuring one capability from two angles. If they invert, we have two capabilities and should stop using the word progress without saying which one we mean.

Sources

  1. Time Horizon 1.1 (METR, Jan 2026)
  2. Announcing ARC-AGI-3 (ARC Prize, Mar 2026)
  3. ARC-AGI-3 benchmark page (ARC Prize)
  4. ARC-AGI-3: The New Interactive Reasoning Benchmark (DataCamp, Mar 2026)