What the original study found

On July 10, 2025, METR published a randomised controlled trial in which 16 experienced open-source developers worked on 246 real issues from their own repositories, with each issue randomly assigned to allow or forbid AI tools. The repositories averaged more than 22,000 stars and over a million lines of code. Tasks took about two hours on average. Developers were paid $150 an hour, recorded their screens, and used Cursor Pro with Claude 3.5 and 3.7 Sonnet when AI was allowed.

When AI was allowed, developers took 19 percent longer to finish. The confidence interval ran from 2 percent to 39 percent slower, so the direction was established even if the size was not. Before the study developers forecast a 24 percent speedup. After finishing, having lived through the slowdown, they still believed AI had made them 20 percent faster. That gap between measured and perceived effect is the finding that travelled furthest, and we think deservedly so.

METR was careful about scope. They described the result as a snapshot of early-2025 capability in one setting, and listed conditions under which it would not transfer: less experienced developers, unfamiliar codebases, and different tools. They investigated around 20 candidate factors and named five as likely contributors. Most of the commentary that cited the 19 percent did not carry the caveats along with it.

The follow-up and what it could not do

On February 24, 2026, METR published the results of a second wave run from August 2025 onward. This time there were 57 developers across 143 repositories and more than 800 tasks, 10 of them returning from the original cohort and 47 newly recruited, paid $50 an hour rather than $150. Tools had moved on to Claude Code and Codex.

Among the returning developers the point estimate was an 18 percent speedup, with a confidence interval from 38 percent faster to 9 percent slower. Among the new developers it was a 4 percent speedup, with an interval from 15 percent faster to 9 percent slower. Both intervals cross zero. METR said the effect had likely flipped positive, and we read the data the same way, but neither number is something we would quote without the interval attached.

The more interesting part of the update is the section on why the design broke. Three things went wrong, and all three are consequences of the tools getting better rather than flaws in the original method.

Selection: the developers who would not participate

The first problem is that the treatment became something people refused to give up. METR reported a significant increase in developers declining to take part, because they would not work without AI, and 30 to 50 percent of participants said they were choosing not to submit some tasks because they did not want to do them without AI tools. That is a selection effect operating at two levels at once. The developers who remain are the ones least attached to the tools, and the tasks they submit are the ones they were willing to do by hand.

Think about what that does to the estimate. If developers withhold exactly the tasks where they expect the biggest uplift from AI, the measured no-AI condition is drawn from the easier end of the distribution and the measured uplift is biased toward zero. You cannot fix this by recruiting more people, because the refusal is a function of the treatment itself. A study of a drug that patients refuse to be randomised away from has the same problem, and the answer there was never a bigger trial.

Measurement: what "time on task" means with agents

The second problem is that the outcome variable stopped being well defined. The original study measured wall-clock implementation time from screen recordings and self-reports, which is a sensible measure when the developer is the one typing. With Claude Code and Codex, METR found that developers would start an agent, then go and work on an unrelated task while it ran. Time on task is now the union of intervals during which either the human or the agent was working, and the two do not have to overlap.

The third problem compounds it. A fraction of developers ran several agents at once on different tasks. METR said measurements were unreliable for that group. A developer who has three agents running and is reviewing the output of one while the others work is spending a third of an hour on each task, or a full hour, or something in between depending on what you are trying to measure. Wall-clock time per task, the thing the original design randomised over, is no longer a property of the task.

A worked example makes the ambiguity concrete. Suppose a developer starts a two-hour task at nine, kicks off an agent at 9:10, switches to a different repository until 9:50, reviews the agent output, fixes two things, and pushes at 10:20. Time on the task is 80 minutes by the clock, 40 minutes by human attention, and somewhere around 110 minutes if you count agent compute time. The original study would have recorded 80. The productivity the developer cares about is closest to 40. The cost the employer pays is closest to the third number.

What we take from the pair of results

The original 19 percent is a real result about a real moment. Early 2025 tools, in the hands of people who knew their codebases very well, on tasks those people had chosen, slowed them down while feeling like a speedup. That should be permanently unsettling for anyone who evaluates tooling by asking users how they feel. The February update suggests the sign flipped as tools became agentic, and we believe the direction, but the honest reading of the intervals is that the magnitude in this population is somewhere between a small slowdown and a large speedup.

The deeper lesson is about what the method can carry. Task-level randomisation with time as the outcome was a good design for a tool that assisted a person. It is a poor design for a tool that works in parallel with a person who has stopped agreeing to work without it. METR listed where they are heading next, which includes observational data, questionnaires, fixed-task experiments, agent evaluations and developer-level randomisation. Of those we would put our money on fixed-task experiments with a throughput outcome, where the question is how much a developer ships per week rather than how long any one task takes.

What we would want someone to try is to run that design and, in the same study, keep asking the perception question. The 24 percent forecast and the 20 percent belief from 2025 are data too. If perception has tracked reality as the tools improved, that tells us something about how developers learn. If it has not, that tells us something more important.

Sources

  1. METR: Measuring the impact of early-2025 AI on experienced open-source developer productivity
  2. METR: Developer productivity uplift update (February 2026)