The 2026 AI Index and the disappearing academic frontier
Stanford's ninth AI Index reports that industry built over ninety percent of notable models last year and disclosed less about them than the year before. This is what a non-profit without a cluster can still do about that.
Three numbers
Stanford HAI released its ninth AI Index this month, and three numbers from it describe the situation of a group like ours better than anything we could write. Industry produced over 90 percent of the notable frontier models of 2025. The average score on the Foundation Model Transparency Index fell from 58 to 40, reversing a climb from 37 in 2023. And the performance gap between the best American and Chinese models, which the Index says has "effectively closed," stood at 2.7 percent in March, with the lead having changed hands several times since early 2025.
The chapter on research and development adds the detail that makes the transparency number concrete. "Training code, parameter counts, dataset sizes, and training duration are no longer disclosed for several of the most resource-intensive systems, including those from OpenAI, Anthropic, and Google." The most capable models are now the least documented. Of 95 notable models released in 2025, 80 came out without training code. Parameter counts, where reported at all, have sat near a trillion for three years, while training compute, which can be estimated from outside, has kept climbing.
We are writing this as someone at an independent non-profit that will never train one of those 95 models. The question we care about is what is left for us, and the Index, read carefully, gives a more useful answer than the headline implies.
What the frontier being private actually removes
It removes the ability to do science on the frontier model as an object. If the training data, the compute budget and the parameter count are undisclosed, then a scaling result, a data-efficiency result or an architecture comparison cannot be done on the systems that matter, only on proxies. The Index notes that "almost all leading frontier AI model developers report results on capability benchmarks, but reporting on responsible AI benchmarks remains spotty." We get the numbers the labs want to publish, on the benchmarks they choose, for models we cannot inspect.
It also removes the human pipeline. The Index reports that the number of AI researchers moving to the United States has dropped 89 percent since 2017, with an 80 percent decline in the last year alone. New AI PhDs in the US and Canada were up 22 percent between 2022 and 2024, and the Index says the new cohort mostly took academic jobs. That is a strange combination. More people are being trained to do research on systems that fewer of them will ever be allowed to see the inside of.
What it does not remove is the ability to do science on model behaviour. The frontier is closed at the level of weights and data. It is wide open at the level of outputs, because everyone can buy a million tokens. And the Index's own most striking findings this year are behavioural findings that needed no privileged access.
The jagged frontier is the opening
The Index reports that a model earned a gold medal at the International Mathematical Olympiad in the same year that the top model read an analog clock correctly 50.1 percent of the time. Agent success on OSWorld went from 12 percent to around 66 percent, which means agents "still fail roughly 1 in 3 attempts on structured benchmarks." SWE-bench Verified went from 60 percent to nearly 100 in a year, which tells you as much about the benchmark as the models. Documented AI incidents rose 55 percent, from 233 to 362. And the Index states flatly that "improving one responsible AI dimension, such as safety, can degrade another, such as accuracy."
Every one of those findings is the kind of thing an independent group can produce, extend or check with an API key and discipline. The clock result is a probe. The OSWorld number is an evaluation harness. The incident count is documentation. The safety-versus-accuracy trade-off is the sort of measurement that labs have an obvious incentive not to publish and that outsiders have an obvious reason to. When the frontier is private, characterising its behaviour becomes the public science. That is where most of the surprises in this year's Index came from.
The compute picture supports this. Global AI compute grew 3.3 times a year since 2022 and now amounts to the equivalent of 17.1 million H100s, with Nvidia supplying over 60 percent. None of that is available to us and none of it is needed for the work we have just described. Inference is cheap. Judgement about what to measure is the scarce input.
Timelines are shortening again, which raises the stakes
The other thing that has changed this spring is the mood about timelines, which had lengthened through 2025 and has tightened again. Rob Wiblin of 80,000 Hours, who had been a sceptic on near-term AGI, now says his own estimates have compressed by about a year and that automated AI research by 2028 is "plausible if today's trends keep marching on." He cites METR's task horizon doubling every four months rather than seven, and internal reports that Claude writes 80 percent of the merged code at Anthropic. He also says the models remain "pretty bad at messy real-world tasks," and offers AI-managed cafés that burned through their operating capital as evidence.
Hold those two observations together. Capability on structured tasks is compounding fast, and performance on unstructured tasks is poor, and the only people with full access to the models are the companies whose valuations depend on the first fact being the story. Someone independent needs to be measuring the second fact carefully, on a schedule, with published methods. The Index does it once a year. That is not often enough if the doubling time is four months.
What we are going to do
Our plan for the year follows from this. We will not compete on models. We will run a small number of behavioural measurements on the frontier models continuously, publish the harness and the raw outputs, and treat the gap between structured and unstructured performance as the quantity of interest. We will document incidents in our own domains rather than waiting for the annual count. And we will spend our compute on the open-weight models, where the Index says the gap to the closed frontier widened to 3.3 percent last year, because those are the only frontier-adjacent systems on which a scaling or data result can still be checked by someone who did not build them.
The frontier being private gives independent research a reason to change subject rather than to stop. If the labs will not say what is in the models, the useful thing to know becomes what the models do, and that is a measurement anyone can make. We would like to see more non-profits make it, and we would like to see next year's Index cite a few of them.
Sources
From the foundation