The experiment

Jessica McFadyen and colleagues at the UK AI Security Institute, with coauthors at Oxford and Harvard, asked a question that benchmark tables quietly assume away. If you give a model far more inference compute than the benchmark's published protocol allows, how much does the score move, and does it move the same way for every model and every task?

Their main design is fully crossed. Six models across two families, Opus 4, 4.5 and 4.6 from Anthropic and GPT-5, 5.2 and 5.4 from OpenAI, all released between May 2025 and March 2026, run on five benchmarks under one shared ReAct-style scaffold in Inspect. The benchmarks are TerminalBench 2.0, SWE-Bench Pro, FrontierMath, HealthBench Hard and Humanity's Last Exam. Two cyber evaluations from earlier AISI work, a 71-task capture-the-flag suite and a single long-horizon range called The Last Ones, bring the model count to twelve including o3, the Codex variants and the Mythos Preview checkpoint. Each task, model and condition cell gets five independent trajectories.

Three interventions are applied uniformly. Per-trajectory token caps of 5 to 30 million, which is one to three orders of magnitude above typical published budgets. Context compaction, where earlier turns are summarised once the running context passes 130k tokens. And iterative resubmission with a cap of 999 submissions, run under two conditions: no feedback beyond an acknowledgement, or an oracle that says whether the submission was correct.

Where the budget mattered

The headline table compares each benchmark's typical published budget with the evaluated cap. For FrontierMath, going from the usual 1 million tokens to 10 million added 11.7 percentage points on average across models, with a standard deviation of 11. For Humanity's Last Exam, going from 64k output tokens to 5 million added 11.9 points, and more under oracle feedback, 15.5, than without it, 8.3. Those are large moves for benchmarks that are used to rank frontier models against each other.

The software engineering benchmarks barely moved. SWE-Bench Pro is already commonly run at up to 16 million tokens, and nearly doubling that to 30 million bought 0.3 points. TerminalBench went from 7.3 million to 10 million and gained about 1.3 points. HealthBench was the odd one out in the other direction. It has a small typical budget of 16k tokens and still showed only 0.3 points of headroom, with every model plateauing inside the typical budget. The authors are careful to say this is a property of their protocol and does not rule out gains under different scaffolds.

What newer models do with the extra budget

The cross-generation result is the one we expect to be cited most. The paper decomposes improvement at the task level into reach, the fraction of tasks a model solves in at least two of five trajectories, efficiency, tokens needed per solve, and reliability, the fraction of trajectories that succeed on a task the model can reach. Later generations show a significant positive reach effect on all six benchmarks analysed, largest on FrontierMath and the cyber CTFs, and the gains are concentrated on harder tasks for four of the six.

Efficiency is the surprise. On most benchmarks later models did not solve reachable tasks with fewer tokens. The improvement came from reaching harder tasks and solving them more consistently, which only shows up if the budget is large enough for the harder tasks to be attempted at all. That is the mechanism behind the paper's second finding: fixed-budget evaluations understate the frontier more as models improve, because the capabilities that separate generations sit in the region the typical budget cuts off.

Serial, parallel and feedback

The third finding is that there is no single intervention that explains the gains. Repeated submission helped on every benchmark. Correctness feedback mattered most where it could steer continued search, on HLE and SWE-Bench Pro. Spreading a fixed budget across many shallow trajectories rather than one deep one, which the authors estimate by reanalysing the same trajectories, helped most on the stateless benchmarks, HLE and HealthBench, and least on the stateful ones where an environment persists across turns.

There is a small detail in a footnote that we think deserves a section of its own in a follow-up. Scoring the no-feedback condition by the model's first or last submission, rather than by the best submission, lowers solve rates on most benchmarks, most strongly on HealthBench, where models reach a correct answer and then regress away from it. A model that cannot tell when it is done is a different object from a model that cannot get there, and the two are indistinguishable under best-of-n scoring.

Why this matters for safety thresholds

The recommendation is to report capability as a curve over inference spend, to state the protocol explicitly, and to compare generations over a large shared compute range at matched budgets. We would go further for safety-relevant evaluations. A threshold like the ones in the published frontier frameworks is usually a number on a benchmark, and this paper shows that the number depends on a budget the evaluator chose. A cyber evaluation that reports 50 percent at the published budget might be at 80 percent at a budget a well-resourced attacker can afford. The CTF curves in this paper were still climbing at 50 million tokens for all ten models.

The question that matters for a threshold is where on the spend axis the model crosses it, and whether that point is within reach of the actors the threshold is meant to worry about. That is a more expensive evaluation to run, and the paper's own caps show what it costs. We think it is the only version that means anything, and we would want the next round of frontier framework reports to plot the curve rather than the point.

Sources

  1. McFadyen et al., How Inference Compute Shapes Frontier LLM Evaluation (arXiv 2606.17930)