The experiment that carries the paper

Sayash Kapoor, Arvind Narayanan and three colleagues at Princeton posted a paper on July 1 that we think every agent benchmark maintainer should read. The central experiment is small. They take HumanEval, the coding benchmark that most agent papers report on, and rerun several published agent architectures while recording the dollar cost of each run alongside accuracy. Then they add some baselines that nobody would write a paper about.

LATS with GPT-4 reaches 88 percent at a cost of $134.50 for the run. LDB reaches 91 percent for $2.19. Reflexion reaches 87.8 percent for $3.90. A baseline the authors call warming, which is GPT-4 with a temperature schedule and retries, reaches 93.2 percent for $2.45. An escalation baseline that starts with a cheap model and moves up only on failure reaches 85 percent for $0.27. The most expensive architecture on the list is over fifty times the price of the baseline that beats it.

Why accuracy alone produces this

The mechanism is not mysterious. Sampling from a language model is stochastic, so calling it more times and picking the best output raises accuracy. Any architecture that makes more calls will climb the leaderboard, whether or not the extra machinery around those calls is doing anything. A leaderboard that reports accuracy and nothing else cannot tell the difference between a good idea and a large sampling budget, and so it rewards both equally.

The practical consequence is that a developer picking an agent from a leaderboard is picking the one that spent the most, and would do as well or better with a retry loop. The authors' framing is that the community has reached mistaken conclusions about which agent designs work, because the evaluations could not have shown otherwise.

Two audiences, one number

The paper's second point is that model developers and downstream developers need different benchmarks and have been sharing one. A model developer wants a cost proxy that is stable over time, and parameter count serves. A downstream developer needs the dollar cost of running the thing in production, which changes with pricing and with how the benchmark is structured.

Their example is NovelQA, a long-context benchmark that loads a novel once and then asks many questions against it. That makes long-context models look cheap relative to retrieval. In real use the questions arrive one at a time and the novel is reprocessed each time, and under that pattern retrieval costs about twenty times less. A downstream developer reading the leaderboard would pick the wrong approach for their workload. The benchmark is accurate about the models and silent about the usage pattern, and the usage pattern is what decides the bill.

No holdout, no generality

The third finding is about overfitting. The authors sort agent benchmarks into four levels of generality, from distribution-specific tasks like math problems up to fully general cross-domain agents, and argue that each level needs a different kind of holdout. A domain-general web agent needs unseen tasks, not just unseen instances. Of eight domain-general benchmarks they examined, one had a task-level holdout. Seven had none.

The worked case is STeP on WebArena, the top-scoring agent at 35.8 percent. Its policies were hardcoded for specific task patterns, including URL templates such as the profile path on the benchmark's Reddit clone. That gets the score. It also means the agent fails the moment the site changes its routes, which is what a holdout set would have caught. The benchmark could not distinguish an agent that browses from one that memorised the benchmark, so it rewarded the second.

There is a list of reproducibility problems in the same section that we recognise from our own attempts to rerun agent papers. Evaluation scripts that assume one agent design. HumanEval as shipped for agents missing test cases for three problems. Costs too high to run enough seeds for a confidence interval. Rate limits on WebArena that break the independence of tasks. And bugs in released code from both LATS and STeP that marked failed tasks as passed.

What to do

The recommendations are straightforward and the hard part is adoption. Report dollar cost next to accuracy and plot the Pareto frontier rather than a single ranking. Report token counts as well, so results remain interpretable after prices change. Build holdout sets matched to the level of generality the benchmark claims. And separate the leaderboards for model developers from the ones for people who ship agents, because a single number serves neither.

The thing we would add is that cost-controlled evaluation needs to start from the baselines the paper used. Retries, warming and escalation cost a few dollars to implement and should be the floor any published architecture has to clear. If an agent paper does not beat a retry loop at equal cost, the architecture has not been shown to do anything, and the reviewers should say so. Our guess is that a large fraction of the agent literature from the past eighteen months would not survive that test, and it would be better to find out now than to keep building on it.

Sources

  1. Kapoor et al., AI Agents That Matter (arXiv 2407.01502)
  2. AI Agents That Matter, full text (ar5iv)