What was actually said

In December 2024 Ilya Sutskever stood up at NeurIPS and said that pre-training as we know it will end. The supporting lines were just as quotable. We have achieved peak data and there will be no more. There is only one internet. Data is the fossil fuel of AI. The room took it as a verdict on scaling, and within a week the phrase had been flattened into a headline that scaling had hit a wall.

That is not quite what he said. The claim was about pre-training on human text, with the observation that compute keeps growing and the internet does not. The Amplify Partners summary of the conference captured the mood accurately. Progress would shift from pre-training scale to inference-time compute, and the log-shaped return on pre-training compute was starting to bite. Sutskever's own answer to what comes next was agents and reasoning, and a comparison to the human brain, which stopped growing in size while humanity kept advancing.

We want to revisit the argument now because the last twelve months have produced evidence on both sides, and because the two camps mostly talk past each other about what the word wall means.

The case that there was no wall

Tomasz Tunguz published a post on November 20 titled The Scaling Wall Was A Mirage, and it is the cleanest statement of the optimist case. Gemini 3, he reports, has roughly the same parameter count as Gemini 2.5, about one trillion, and yet became the first model to break 1500 Elo on LMArena and beat GPT-5.1 on 19 of 20 benchmarks. He quotes Oriol Vinyals of Google DeepMind attributing the gains to improving pre-training and post-training, and saying the delta between 2.5 and 3.0 is as big as Google has ever seen with no walls in sight.

The second half of Tunguz's argument is about hardware and money. Nvidia's data center revenue in the quarter was 51 billion dollars, up 66% year over year, and the company projects three to four trillion dollars of annual AI infrastructure spend by the end of the decade. The inference is that the people with the best information are still buying, so the returns must still be there.

The Vinyals quote is the important part and it is worth reading carefully. It credits pre-training and post-training together. It does not say the gains came from more tokens of human text. A model with the same parameter count getting much better is consistent with better data, better post-training, and more RL, and it is also consistent with the thing Sutskever said would happen after pre-training as we knew it ended.

The case that the wall is real and we walked around it

The data argument has not gone away. Epoch AI's 2024 estimate put the stock of public human text at around 300 trillion tokens with exhaustion projected between 2026 and 2032, sooner for overtrained models. Nothing in 2025 changed that arithmetic. What changed is where labs get their training signal. Reasoning models trained with reinforcement learning on verifiable problems generate their own data. Synthetic corpora fill in for scarce domains. Post-training budgets that used to be a rounding error are now a substantial fraction of the total.

So both sides can point at Gemini 3 and be satisfied. The optimist sees a huge capability jump at constant parameter count and says scaling is fine. The pessimist sees a jump that came from everything except more human text and says pre-training as we knew it did end, quietly, and got replaced. The disagreement is mostly about the definition of scaling, and we think the pessimist has the more precise version.

Why we distrust both narratives

The evidence in this debate is thin in a specific way. Nobody outside the frontier labs knows the token counts, the data mix, or the split of compute between pre-training and RL for any 2025 model. The parameter count Tunguz cites for Gemini 3 is reported rather than disclosed, and the attribution of gains is a single sentence from an executive. Investor spending tells you what people expect, which is a different quantity from what is true.

From the outside, the strongest thing we can say is that capability per parameter went up a lot in 2025 and that the labs credit a mix of methods. That is compatible with pre-training scaling laws being intact and compatible with them having flattened while other levers took over. A clean test would hold the post-training recipe fixed and vary only pre-training compute on a modern data mix. Nobody with the budget to run that test has any reason to publish it.

What we think comes next

Our reading is that Sutskever was right about the constraint and wrong, or at least early, about the consequence. The internet did not grow. The cost of that turned out to be small, because the industry had two years of warning and used it to build data sources that do not depend on the internet growing. If there is a wall in 2026 it will be a verifier wall rather than a data wall, meaning the supply of tasks where a model can check its own work runs out before the compute does.

The experiment we would fund is small and public. Take an open base model, fix a post-training recipe, and pre-train three versions at 1x, 4x and 16x compute on a documented 2025-style data mix that includes synthetic and RL-generated text. Then measure whether the post-training gains stack on top of the pre-training gains or substitute for them. That would tell us more about the wall than any number of benchmark tables from models we cannot inspect.

Sources

  1. NeurIPS 2024: main themes and takeaways (Amplify Partners)
  2. Techmeme coverage of Sutskever's NeurIPS 2024 talk
  3. The scaling wall was a mirage (Tomasz Tunguz)
  4. Will we run out of data? (Epoch AI, 2024)