The numbers as reported

Stanford HAI published this year's AI Index this month. The figure we keep returning to is the count of notable machine learning models by originating sector. In 2023 industry produced 51, academia produced 15, and industry-academia collaborations produced 21, which the report calls a new high for that category. By country the United States accounted for 61 notable models, the European Union 21, and China 15.

The report also puts prices on the frontier. Its estimate for GPT-4's training compute is 78 million dollars. For Gemini Ultra it is 191 million dollars. These are estimates built from public information about hardware and training duration, so treat them as order-of-magnitude figures rather than invoices, but the order of magnitude is the point.

One more figure from the same report. Funding for generative AI nearly octupled from 2022 to reach 25.2 billion dollars in 2023. Whatever the exact numbers, the money and the models are concentrating in the same place.

What academia can no longer do

A university group cannot train a GPT-4. We do not think anyone disputes this, but it is worth being precise about what follows. The ratio of 51 to 15 does not mean academic labs got worse or less ambitious. It means the definition of a notable model moved to a scale that only a handful of companies can pay for, and the count follows the definition.

The consequence for research is that the most capable systems are studied mainly by the people who built them, under whatever disclosure rules those people set. The Index says this plainly in a different section. OpenAI, Google, and Anthropic each test their models primarily against different responsible AI benchmarks, so there is no consistent way to compare the risks and limitations of the leading systems from outside. That is a property of the field's structure, and the sector count is the structure.

Here is a worked example of the constraint. Suppose a graduate student wants to test whether a training intervention changes how a model generalises. At the scale of the models in the 15, she can train the model herself, vary the intervention, and report the result. At the scale of the models in the 51, she can do that only if a company gives her access, which they may, on their terms, to their checkpoints, with their training data hidden. The question is the same. Only one version of it is science she controls.

What academia can still do, and does

The 21 collaborations are the encouraging line, and we read them as the shape of the next few years. A company brings compute and a trained model. A university brings people who will ask questions the company did not think to ask, and who will publish the answer whether or not it flatters the sponsor. That last property is the one worth protecting in any collaboration agreement.

The other thing academia does well is study small things carefully. Most of what we know about how transformers learn came from models far below the frontier, and the Index's own performance tables show that on competition-level mathematics, visual commonsense reasoning, and planning, the frontier models still trail humans. Those are the gaps where a cheap model and a good experimental design can still produce a result that matters.

There is a third path, which is treating the frontier models as a natural phenomenon and studying them from the outside. Behavioural evaluation, black-box probing, and careful measurement of failure cases need an API key rather than a datacentre. The limitation is that you cannot tell what changed between two versions, and the company can retire the version you studied without notice. Reproducibility becomes something you hope for rather than something you build.

What we think the number is for

The Index reports the sector split every year and it will get more lopsided before it gets less. We do not think the right response is to fund universities to compete at 78 million dollars per model. The right response is to be clear-eyed about what academic AI research is now, which is the study of systems built by others, plus the study of small systems that stand in for them, plus the construction of instruments that let outsiders measure the large ones.

The instruments are the piece that most needs work. If the three leading developers evaluate against three different benchmark sets, then a shared, independently run evaluation is worth more to the field than another 7 billion parameter model would be. That is a contribution a university can make at university prices, and it is the one the sector count says we most lack.

The share of Americans who say they are more concerned than excited about AI rose from 38 to 52 percent in the report's public opinion section, which tells us people want someone independent checking these systems. Fifteen models a year is a small share of the frontier. Fifteen good evaluations a year would be a large share of what the public actually needs.

Sources

  1. Stanford HAI, Artificial Intelligence Index Report 2024