The reframing

Greg Burnham, who heads the benchmarks work at Epoch AI, published a piece on August 14 that reads as a design brief for anyone about to build an evaluation. His argument is that the usual question, can the model do this task, is the wrong unit. A benchmark is worth building when the score would change what you believe about a larger question, and he lists nine of those questions. We have spent a good part of the last year arguing with colleagues about which evals matter, and this is the cleanest framing of the argument we have read.

The reason we find it useful is that it gives a rejection criterion. If you cannot say which of the big questions a new benchmark would move, then the benchmark is probably measuring something that a model card will report and nobody will act on. Below we go through the nine as Burnham states them and add what we think each one demands of the evaluation design.

Jobs, economic value and the frontier gap

The first question is whether AI can do a whole job rather than pieces of one. Burnham points to polling showing people still use AI for parts of tasks, and names MirrorCode on large software packages, the Remote Labor Index on freelance projects, and Andon Café, an AI-operated real cafe, as attempts at whole-job measurement. The design demand here is length and messiness. A job has ambiguous specs, interruptions and long horizons, and a benchmark that cleans those away is measuring a task again.

The second is whether progress is landing in economically impactful areas, and he lists cybersecurity, computer use and guidance for physical industry as examples. The third is how consistent the gap between frontier and trailing models is across domains, which he ties to whether leading developers can capture enough value to justify their infrastructure spending, and to the open versus closed weights split between American and Chinese labs. He raises benchmaxxing as the obvious confound. A gap measured on a benchmark that trailing labs target directly tells you about targeting, and a gap on a held-out domain tells you about capability.

Why scores correlate and whether AI can do AI research

The fourth question is one we had not seen stated so plainly. Benchmark scores are all correlated, and Epoch's own Capabilities Index leans on that correlation to combine many evaluations into one number. Burnham asks whether the correlation reflects a real general capability or something else, and notes that METR's time horizon measurement tracked the index before saturating. If the correlation is an artefact of shared training data or shared contamination, then every composite index is measuring the artefact.

The fifth is whether AI can do AI research and development, which matters because it is the input to any recursive improvement story. He is candid about the difficulty. Frontier labs are opaque, so a realistic task is hard to construct from outside, and the compute cost of running an R&D evaluation is high. We would add that this is the question where a benchmark with a single score is least useful, because the thing you want to know is which parts of the research loop have been automated and which have not.

Learning on the fly, inference scaling and RL generalisation

Question six is whether models can learn during deployment through context management alone, without weight updates. Burnham's reason for caring is a safety one. If a model can meaningfully improve after pre-deployment testing, then pre-deployment testing loses much of its force. His named example is EBR-bench, built on repeated campaign-style board games, and he reports that models have so far failed to show much improvement across repeated play. That is a rare case where a null result on a benchmark is the informative outcome.

Question seven is the returns to inference scaling. The bigger question underneath is whether enough thinking time lets a model solve anything, and at what cost. He notes that nobody has settled what the optimal inference scaling strategy is and that the answer may differ by domain, which means a single scaling curve on one benchmark is not going to resolve it. Question eight is how well reinforcement learning generalises across domains and out of distribution, and here his evidence is positive. Strong results on undisclosed game puzzles and on programming in low-resource languages suggest the flexible reasoning story has something to it, rather than the story that training data quality does all the work.

New ideas, and what we would do with this list

The ninth question is whether AI can produce new ideas. Burnham's example is FrontierMath: Open Problems, which is built so that a solution would require an insight that is not already in the literature. He also adds the caveat we would have added, which is that even a model that generates new ideas is bottlenecked by physical testing in most fields, so the real-world effect may lag the benchmark by a long way.

Reading the list against the benchmarks we have seen over the last two years, most of them speak to questions one, three and seven and almost none to four, six and nine. That is partly because the easy questions are easy to instrument, and partly because the hard ones need a null result to be interesting, and nobody gets credit for a benchmark on which every model scores zero. Our suggestion is small. When you propose a new evaluation, write down which of these nine it would move and what score would change your mind. If the answer is none of them, build something else.

Sources

  1. Greg Burnham, 9 big questions benchmarks can help answer (Epoch AI Gradient Updates)