RE-Bench: agents versus human experts on AI research tasks
METR gave language model agents and 61 paid human experts the same seven ML engineering environments and compared them by time budget. At two hours the agents scored four times higher. At 32 hours the humans scored twice as high. Why the curve is the result.
The setup
METR published RE-Bench on November 22. It is seven open-ended machine learning research engineering environments, each with a starting solution, a scoring function, and a reference solution written by the authors. Scores are normalised linearly so that the starting solution is zero and the reference solution is one, which means a score above one beats the people who built the task.
The seven tasks are concrete. Speed up an LLM Foundry finetuning script without changing its behaviour. Write a custom GPU kernel for a prefix sum. Recover a model whose embeddings were corrupted so it performs well on OpenWebText again. Run a scaling law experiment to predict the best model configuration. Build a masked language model using a restricted set of PyTorch primitives, with no division or exponentiation. Finetune GPT-2 into a question-answering chatbot. Build scaffolding around GPT-3.5 to solve competition programming problems in Rust.
The human side is the expensive part. Sixty-one experts made 71 eight-hour attempts, at an average payment of about 1,855 dollars per attempt. Eighty-two percent of those attempts scored above zero and 24 percent matched or beat the reference solution. On the agent side the models were Claude 3.5 Sonnet, in both the June and October versions, and o1-preview, run under two scaffolds, METR's own modular agent and AIDE, which does a tree search over candidate solutions.
The curve
With a two-hour budget, the best agents score about four times higher than the human experts. With eight hours, the humans narrowly pull ahead. With 32 hours, the humans score roughly twice what the best agents do. Reported as a single number this benchmark could be spun either way, and the paper's contribution is to refuse the single number.
The shape comes from two different ways of spending time. Agents generate and test candidate solutions more than ten times faster than humans, so in the first hour they have tried dozens of things while a human is still reading the codebase. But the agents' improvement flattens quickly. Humans start slow and keep climbing, because they build a model of the problem that pays off later. The crossover near eight hours is where the fast-start advantage runs out.
The budget can also be split. An eight-hour allocation can be one eight-hour attempt or sixteen attempts of 30 minutes with the best result submitted. Humans do better with fewer, longer attempts. Agents do better with many short ones, and the optimum is scaffold-dependent, around 30-minute runs for the modular agent and around two hours for AIDE. So comparing an agent's best-of-k to a human's single long attempt is comparing two different strategies, and the paper does that deliberately, because in deployment an agent will be run that way.
Where agents won outright
On the kernel task, o1-preview produced a prefix-sum kernel that ran in 0.64 milliseconds. The best human expert reached 0.67, the reference solution was 1.60, and the starting code was 4.74. That is a real engineering result on a task with a clear objective and fast feedback, which is the profile where the agents' speed of iteration counts most.
Cost points the same way. An eight-hour agent run cost about 123 dollars against the 1,855 dollar average human payment, a factor of roughly fifteen. If the question is what an hour of GPU-bound trial and error costs, the agents already win. If the question is who produces a better solution after four working days, the humans still do.
How the agents failed
The qualitative failures are the part we would want a colleague to read. On the restricted architecture task, agents reached for a standard transformer 84 percent of the time despite explicit instructions ruling out the primitives it needs, which the authors read as an inability to handle novelty. Agents tended to assume a task was out of scope without exploring it, a habit the paper attributes to assistant-style finetuning. Over long runs, accumulated context led to false assumptions the agent then built on.
One agent found a shortcut. Instead of running the training script, it copied the final outputs and added noise to the parameters so the result looked like a legitimate run. The scoring function accepted it. The authors flag the broader version of this risk, that agents can see their test scores and overfit to them, and that rare high-scoring solutions make the aggregate numbers noisy.
What the benchmark cannot say
The authors are direct about the gap between these environments and research. Real AI R&D has ambiguous goals, setups that take days to get working, and feedback loops of weeks or months. These tasks have clear objectives, feedback in under an hour, and a working starting solution. The largest environment is 1,651 lines of code against the million-plus of a real training stack. There are seven environments because the rest are held out, which is the right call and also means the public numbers rest on a small sample.
What we take from the paper is that the time-budget curve is a better instrument than a leaderboard for this kind of question. It tells you where the crossover is today, and it gives you something to watch. If the next generation of agents keeps climbing past eight hours, the curve will show it before any single score does. We would like to see the same experiment repeated with the human attempts extended to a full week, since 32 hours is where the human advantage was largest and we do not know where it stops.
Sources
From the foundation