The measurement

David Gringras and Misha Salahshoor posted an audit on May 5 that puts a number on something most of us have suspected. They pulled 112,303 records matching LLM keywords from January 2022 to April 2026, extracted 4,766 full texts, and for each paper worked out which model was evaluated and how capable the frontier was at the time. The capability scale is the Epoch Capabilities Index, anchored at 150 for GPT-5 in August 2025 and calibrated across roughly 165 models using 1,471 benchmark entries.

The median paper evaluated a model 10.85 ECI points behind the contemporaneous frontier. To make that concrete, the authors say it is about 1.4 times the distance between Claude Sonnet 3.7 and Claude Opus 4.5, so the typical paper is testing something a full major version behind what was available. The gap is also growing, at 5.53 ECI points a year with a 95 percent interval of 5.03 to 5.83. Papers are falling further behind, not catching up.

The reporting gap is worse than the lag

The lag has innocent explanations. Peer review takes months, API access costs money, and a paper submitted in January is not going to evaluate a model released in March. The authors say as much, describing the pattern as the predictable balance of review time, API cost and reporting standards written before reasoning models existed. They are careful to frame it as a structural problem rather than misconduct.

What has no innocent explanation is the reporting. Of abstracts evaluating models with a reasoning mode, 3.2 percent said whether it was on. In full texts the figure is 21.2 percent. Since reasoning mode can move a score by more than a model generation, a paper that omits it has left out the single largest source of variance in its own result. A reader cannot tell whether the finding is about the model or about a cheap configuration of it.

Then there is the framing. In 52.5 percent of abstracts, with an interval of 48.2 to 56.9, conclusions were stated at the level of AI rather than the specific model tested, and that habit is rising with an odds ratio of 1.23 per year. Put the three numbers together and you get the typical paper in the corpus. It tests a model a generation old, in an unreported configuration, and reports the result as a property of AI.

Where the papers come from

The 18,574 papers in the domain analysis span medicine, law, coding, education and scientific reasoning, with medicine the largest share by a wide margin and the other four in single digits to mid teens. About 89 percent concentrate on four model families, OpenAI GPT, Anthropic Claude, Google Gemini and Meta Llama. That concentration matters for the lag finding because these are exactly the families where a new version arrives every few months and where reasoning modes exist to be reported or not.

Medicine dominating the corpus is the part that worries us most. A clinical paper that finds a model unreliable at some task, tested on a version two generations old with reasoning off, will circulate in a field that does not track model releases, and the finding will be read as settled. The authors list several limitations of their own, including keyword sampling bias and the fact that they cannot say whether any given result would survive re-execution on the frontier. That last point is the important one. The audit measures the gap, not what the gap hides.

VERSIO-AI

The proposed fix is a 13 item reporting checklist called VERSIO-AI, covering what the authors call the elicitation surface. That means the model snapshot, the evaluation date, access tier, reasoning mode and effort, tool access, scaffolding, prompting protocol and sampling settings. The full list is in the appendix and is meant to slot into existing frameworks like CONSORT-AI and TRIPOD-LLM rather than replace them.

Three items are proposed as desk-reject criteria, meaning an editor should bounce the paper without review if they are missing. Those are the model identifier, the declared evaluation frame, and the reasoning mode status. We think this is the right cut. None of the three costs anything to report, all three are known to the authors at submission, and their absence is the difference between a result someone can check and one nobody can.

What we would add

The authors direct their recommendations at three groups. Authors should report the elicitation surface. Editors and reviewers should enforce it. Funders should condition grants on disclosure and on API budgets that let a lab test at the frontier rather than at whatever tier the department could afford. The funder recommendation is the one we have not seen elsewhere, and it addresses the actual cause of the lag rather than the symptom.

The thing we would add is a norm on the claims side. If a paper tests one snapshot of one model, the abstract should name it and stop there. The word AI in a conclusion should require evidence across families and across time. That is a norm reviewers could enforce tomorrow, and the audit gives them the number to point at when they do.

Sources

  1. Gringras and Salahshoor, Frontier Lag: A Bibliometric Audit (arXiv 2605.04135)
  2. Full text of the paper (arXiv HTML)