The seven-month doubling: reading METR's time-horizon paper carefully
METR's new metric says frontier models can complete tasks that take humans about an hour, and that the number has doubled every seven months since 2019. Notes on how the number is built and where it is soft.
What the metric is
METR published a paper last week from Thomas Kwa, Ben West, Joel Becker and about two dozen coauthors that proposes a single number for how capable a model is at long tasks. Take a set of tasks. Measure how long skilled humans take to do each one. Fit a curve of model success probability against human time. The 50 percent time horizon is the human duration at which the model succeeds half the time.
The claim that got the attention is the trend. Plotting the horizon for frontier models from GPT-2 in 2019 to the present, they get an exponential with a doubling time of roughly seven months. GPT-2 sits at about two seconds. Claude 3.7 Sonnet, the current frontier, sits at roughly 50 minutes in the paper and about an hour in the blog post. Extrapolate and you get models handling week-long tasks in two to four years, and month-long tasks within about five.
We like the metric more than most single-number capability measures because the unit is legible. Minutes of human work means something to anyone who has done the work. What we want to do here is walk through how the number is assembled, since the assembly is where the assumptions live.
Where the human times come from
The task set is 170 tasks from three suites. HCAST is 97 tasks in 46 families covering software and research work. RE-Bench is 7 machine learning research tasks, each budgeted at eight hours. SWAA is 66 tasks the authors wrote for this paper, each a single action taking under a minute, added so the curve has anchors at the short end where GPT-2 era models live. Human baselines exist for 148 of the 169 tasks that were scored.
The baseline data is substantial. More than 800 people contributed 2,529 hours. For HCAST and RE-Bench there were 558 baseline attempts, of which 286 succeeded. For SWAA there were 249 attempts and 236 successes. That success ratio matters. The human time for a task is computed from successful attempts, so the humans who set the clock are the ones who finished. A model is scored on its success rate against that clock. Those are different quantities and the paper is open about it.
The other thing to hold onto is that humans on HCAST and RE-Bench were given up to eight hours and were unfamiliar with the specific codebases. A task that takes a domain expert who knows the repository twenty minutes may take a contractor who has never seen it two hours, and it is the two hours that enters the metric.
How wide the intervals are
The headline seven months comes with a 95 percent confidence interval from a hierarchical bootstrap over task families, tasks and attempts. In the version we read the interval on the doubling time runs from about 166 to 240 days, roughly plus or minus 19 percent. The blog post is looser and says the trend is consistent with one to four doublings per year depending on which models and tasks you include. Both are honest. The narrow interval is the statistical uncertainty given the data. The wide one is the uncertainty from choosing the data.
The recent segment is steeper. Fit only 2024 and 2025 models and the doubling time drops to about four months. The paper does not commit to which trend is real, and gives extrapolations under both. Under the seven-month trend, a one-month horizon, meaning 167 hours of human work, arrives between mid-2028 and mid-2030 at 80 percent confidence. Under the faster trend it arrives in early 2027. That is a three-year spread in the arrival date of the same milestone from the same dataset.
The 80 percent horizon is the number we would put next to the 50 percent one in any summary. It is about five times shorter. A model that gets half of hour-long tasks right gets four fifths of roughly twelve-minute tasks right. For anyone deciding whether to hand a model a task unsupervised, the second number is the relevant one.
Messiness
The paper includes a messiness score for each task, a rating of how much it resembles real work, with factors like ambiguous success criteria, lack of automatic feedback, irreversible mistakes and unwritten conventions. Models do worse on messier tasks. The trend of improvement is similar across messiness levels, which the authors read as reassuring. We read it differently. If the messy tasks in the set are still clean enough to have a programmatic grader, then the set has a ceiling on messiness, and the trend across that range says little about tasks above it.
That points at the deepest assumption in the metric. Every task has an automatic check for success. That is what makes it possible to run thousands of model attempts. It also means the metric can only see work that can be verified by a test. A lot of real software work is deciding what the test should be, and no task here measures that.
The objections we expect
The objections we expect to see within weeks cluster into four, and we share most of them. Task selection, since 170 tasks with test oracles are not a sample of software work. Baseline quality, since the humans were unfamiliar with the codebases and only their successful attempts count. The asymmetry between a human under an eight-hour limit and a model scored on success rate with no equivalent clock. And the extrapolation itself, which projects a fitted curve years past its data. Put bluntly, if a task has a test oracle, a model will soon pass it half the time, and that says less about competence than the trend line implies.
None of these is fatal to the metric and all of them are fatal to reading the extrapolation as a forecast. The paper's own limitations section makes most of the same points, which is to its credit. Our suggestion is to keep the 50 percent horizon as a way to track progress on the tasks that have oracles, report the 80 percent horizon alongside it, and treat any date past the data as a hypothesis to be tested by the next model rather than a prediction to be planned around. The best thing about a metric with a clear construction is that a critic can say exactly which step they doubt.
Sources
From the foundation