The question and why it is hard

Ask whether AI progress has sped up and you will get a chart with a line through it. The problem is that any noisy upward series can be fit with a line, a curve, or a line with a kink, and each fit tells a different story. Fit a straight line and progress is steady. Fit an exponential and it is accelerating. Fit two lines with a break in late 2024 and reasoning models changed everything. The chart cannot tell you which is right because all three go through the dots.

Jean-Stanislas Denain and Alexander Barry at Epoch AI published a study on April 16 that treats this as a forecasting problem rather than a curve-fitting one. The idea is simple. A model of the trend that is true should predict the future better than one that is false. So instead of asking which curve fits the data best, ask which curve, fit on data up to some date, best predicts the data six months later.

Four metrics, eight models

The four capability series are Epoch's own Capabilities Index, which has data from March 2022, the METR 50 percent time horizon in log form, a combined maths index built from MATH Level 5, FrontierMath, OTIS Mock AIME and MathArena, and a WeirdML V2 index built from per-question accuracy. Together they span general capability, agentic task length, mathematics and a deliberately odd machine-learning coding benchmark.

The eight candidate trends, ordered from slowest to fastest implied growth, are a global linear fit, a reasoning split with separate lines for reasoning and non-reasoning models, a piecewise linear fit with a break, a linear fit with a log term, a quadratic, a power law, an exponential and a hyperbolic curve. That ordering matters because it turns the question from acceleration yes or no into a ladder, and the result is where on the ladder the best predictor sits.

Each series is also prepared four ways, as state of the art on release, as named releases only, as a daily maximum, and as a daily interpolation, so that the conclusion does not hinge on one choice about how to handle the gaps between model launches.

Expanding-window cross-validation

The evaluation is expanding-window cross-validation. Take a minimum training window, fit all eight models on the data inside it, predict the value six months past the window's end, and score the error. Then extend the window forward, refit, and repeat. The minimum windows start between June 2024 and January 2025, so every model gets tested on many different cutoff dates. The primary score is error at the six-month horizon, averaged across cutoffs.

This is the right way to do it and we want to say why. In-sample fit rewards flexibility, so the hyperbolic curve or the quadratic will always look best on the data they were fit to. Out-of-sample prediction penalises flexibility that does not generalise. A model that says acceleration and then over-predicts the next six months loses to a model that says steady and gets the next six months right. The method has a built-in scepticism that a chart does not.

What won

On three of the four metrics, the Capabilities Index, the METR horizon and the maths index, the reasoning split model predicted best. That model says two things at once. Reasoning models arrived with a one-off jump in level, and they also sit on a trend two to three times steeper than the non-reasoning trend. Neither a single straight line nor a smooth exponential predicts as well as that pair of lines.

On WeirdML V2, the global linear fit won. No acceleration, no break, no separate reasoning trend. The authors' suggested explanation is that WeirdML is a resource-constrained environment that has not been a target of reinforcement learning post-training, so the thing that produced the jump elsewhere has not been pointed at it.

What the odd one out tells you

The WeirdML result is the most useful finding in the paper, and the authors flag the reason themselves. The three metrics where acceleration shows up are concentrated in programming and mathematics, exactly the domains where correctness can be checked automatically and where labs have explicitly aimed their reinforcement learning. The metric that did not accelerate is the one that RL has not been optimised for.

So the honest summary is narrower than acceleration. Capabilities accelerated where labs could build a verifier and train against it. Where they could not, or did not, the old trend continued. That is consistent with a world where reasoning training is a powerful but targeted tool, and inconsistent with a world where models got generally smarter at a faster rate across the board. Which world we are in matters a great deal for forecasting, and four metrics are not enough to say.

What we would want next is this exact method applied to a dozen more series, chosen to include several that are hard to verify automatically, like long-form writing quality or open-ended research tasks. If those all look like WeirdML, the acceleration story is a story about verifiers. If some of them look like the maths index, something more general is going on. Either way the cross-validation framework is the thing to keep, and we hope Epoch releases the code.

Sources

  1. Denain and Barry, Epoch AI, Have AI capabilities accelerated? (April 16, 2026)