A forecast with its parameters showing

The AI 2027 scenario was published in April 2025 with a supporting timelines document by Eli Lifland, Nikola Jurkovic and FutureSearch, and this week its authors published a long clarification of how their forecasts have moved since. We have read both closely, because the timelines document is one of the few forecasts in this field that shows its working, and the clarification is one of the few cases of forecasters publicly revising in the open. Both are worth taking seriously, and both are more revealing than the scenario that made them famous.

The forecast targets a milestone the authors call the superhuman coder, defined as a system that would let a company run, on five percent of its compute budget, thirty times as many agents as it has engineers, each working thirty times faster than its best engineer. Two methods are used to estimate when that arrives. One extrapolates METR's time horizon measurements. The other forecasts when RE-Bench saturates and then adds a chain of estimated gaps that must be crossed after that.

The document is unusually candid about its own foundations. Its stated limitation is that "this forecast relies substantially on intuitive judgment, and involves high levels of uncertainty." We want to give credit for that sentence before we spend the rest of this note showing what it means in practice.

What the parameters actually say

Look at one input. The time horizon method needs an estimate of how long a task the superhuman coder must be able to complete. Lifland's estimate is ten years, with an 80 percent interval from one month to 1,200 years. Jurkovic's estimate is a month and a half, with an interval from sixteen hours to 4,000 hours. Those are the two lead authors on the same document, and their central estimates differ by roughly a factor of eighty. The interval on one of them spans four orders of magnitude.

Other parameters are tighter but still wide. The doubling time of the time horizon is put at 4.5 months with a range of 2.5 to 9. The probability that progress is superexponential rather than exponential is set at 0.45 by one author and 0.4 by the other. The gap between RE-Bench saturation and the milestone is decomposed into seven pieces, each with its own median and interval, and the largest of them, the time horizon gap, is 18 months for Lifland with an interval of 2 to 144.

When you push those distributions through the model, the resulting medians are what they are. In April 2025 the time horizon method gave both authors 2027. The benchmarks-and-gaps method gave Lifland 2028 and FutureSearch 2032. By May, Lifland's own update had moved his time horizon result to 2029 and his gaps result to 2030. The all-things-considered medians across the three parties were 2028, 2030 and 2033, with lower bounds in 2026 and upper bounds past 2050. The single year in the scenario's title was never the output of this document. It was one author's modal year, and the document already disagreed with it.

The year of moving medians

The clarification post, by Lifland, Daniel Kokotajlo and Brendan Halstead, lays out the history in a table, and the table is the most useful thing the project has published. Kokotajlo's median for AGI sat at 2027 from December 2022 through January 2025. It moved to 2028 in February 2025, to the end of 2029 in August, to 2030 in November, and to December 2030 this month. Lifland's ran the other way for years, from 2060 in 2021 down to 2031 in April 2025, and has since drifted back out to around 2035.

The reasons given are ordinary. Kokotajlo says the benchmarks-and-gaps analysis convinced him progress was slower than he had assumed, that pretraining appeared to be slowing, and that "things seem to be going somewhat slower than the AI 2027 scenario." Lifland cites intermediate results from a newer model and corrections to the old one. None of this is scandalous. It is what updating looks like.

What is awkward is the authors' complaint that the press has been comparing an "old modal prediction to new model's prediction with median parameters," which they call methodologically flawed. They are right that mode and median are different statistics. They are also the people who put the mode in the title. "We've done our best to make it clear that it has never been the case that we were confident AGI would arrive in 2027," they write, and we believe them, and we also think a document called AI 2027 has forfeited some standing to object when readers take the number literally.

Science or persuasion

Here is our honest reading. The timelines document is a piece of quantitative reasoning with wide, disclosed uncertainty and inputs that are mostly informed guesses. That is a legitimate thing to publish, and it is more than most people making confident claims in either direction have done. The scenario built on top of it is a piece of persuasion. It took a distribution whose 80 percent interval ran from next year to after 2050 and compressed it into a narrative with dates, names and a plot. The narrative was read by far more people than the parameter tables, and it was read as a prediction.

The parameter sensitivity is the real finding. If two co-authors can disagree by a factor of eighty on a key input and the interval on another spans four orders of magnitude, then the priors, rather than the evidence, are what determine the median the model produces, and the priors are what moved during 2025. The model did not learn that AGI was further away. Its authors decided it was, and adjusted the inputs to say so, which is what one should do, and which is also not what a forecast is normally understood to mean.

A forecast done as science would have been retired the way a failed experiment is, with a short statement of what was predicted, what was observed, and which assumption broke. Instead the retirement happened through a series of median adjustments spread over a year, each individually reasonable, and a clarification post arguing that the original number was never really the number. We do not think anyone involved was being dishonest. We think the format made an honest retirement very hard to perform.

What we would want from the next one

Publish the parameter tables first and the story second, and put the interval in the title if you are going to put a year in it. Register in advance which observations would move the median by more than a year, so the moves can be checked against the register when they happen. And separate the modelling team from the storytelling team, since the model here was honest and the story was the thing that got quoted.

We would also like someone outside the project to rerun the model with independently elicited parameters from people who have never read the scenario. If the medians land in the same place, the model is doing work. If they land wherever the new priors put them, that tells us the model is a calculator for beliefs, which is useful, and which should be described as such.

Sources

  1. AI 2027, "Timelines Forecast" (Lifland, Jurkovic, FutureSearch)
  2. Lifland, Kokotajlo and Halstead, "Clarifying how our AI timelines forecasts have changed since AI 2027" (LessWrong)