What he actually said

Sutskever has said very little in public since leaving OpenAI and founding Safe Superintelligence, so a two-hour conversation with Dwarkesh Patel is an event. The framing he offers is a periodisation. From 2012 to 2020 was an age of research. From 2020 to 2025 was an age of scaling, when the recipe was known and the job was to apply more compute to it. Now, in his words, it is back to the age of research again, just with big computers.

The reasons he gives are two. Pre-training data is finite, and he says so without hedging. And he doubts the marginal returns: when Patel presses him, he asks whether anyone really believes that a hundred times more compute would make everything different, and answers that he does not think so. This is a stronger position than the usual talk of diminishing returns. It says the current recipe has a ceiling and the ceiling is close.

The crux is generalisation

The technical claim underneath the periodisation is about generalisation. Sutskever says models generalise dramatically worse than people, and he calls this the crux. A person with fifteen years of experience has seen a tiny fraction of a model's pre-training data and knows far less in total, but knows what they know more deeply. He does not claim to know why, though the interview circles it for a while, and he treats it as the open problem of the next period.

The example he gives of the symptom is familiar to anyone who uses coding assistants. A model does well on evaluations and then, in deployment, fixes one bug and introduces another when asked to correct it. His diagnosis is sharper than the usual one about benchmarks being gamed. He says the real reward hacking is the human researchers who are too focused on the evals. Labs build new reinforcement learning environments aimed at the evaluations they care about, and the models get good at those environments rather than at the underlying skill. That is an argument that the measurement problem and the generalisation problem are the same problem.

Why this is good news for a small lab

If the age of scaling is over, the competitive logic changes. During scaling, the winner is whoever has the most compute, and there is no research question a small group can answer that a large one cannot answer faster. During research, the winner is whoever has the right idea, and ideas do not scale with the size of the cluster. That is the situation in which a small foundation can matter, and it is the situation Sutskever is describing.

He makes the point about his own lab, which has raised three billion dollars and is small by the standards of the companies it competes with. His argument is that the large companies spend most of their compute on inference and most of their money on engineering and product, so the gap in research compute is much smaller than the headline gap. When Patel cites an estimate that one competitor spends five to six billion dollars a year on research experiments alone, Sutskever's reply is that the question is what you do with it, and that there is a lot more demand on training compute at a big lab, with many more work streams. He says SSI has enough compute to prove to itself and to anyone else that what it is doing is correct.

What research means at a million dollars an experiment

Here is where we part from the cheerful reading. The age of research from 2012 to 2020 was one where a good idea could be tested on a single GPU in an afternoon. The age of research Sutskever is announcing has, by his own phrase, big computers in it. An experiment on the generalisation question at anything like frontier scale costs millions. A lab with three billion dollars can run a few hundred of those. A lab with a few million can run none.

So the honest version of the good news is narrower. The next period will reward ideas, but only ideas that can be validated at small scale and then survive the transfer to large scale. That puts a premium on exactly the kind of work small labs already do well: careful ablations, scaling studies on the cheap end of the curve, negative results, and methods that make the question smaller. What it does not reward is small labs trying to imitate the large ones at a hundredth of the budget.

Sutskever's timeline for a system that learns as efficiently as a person and then goes past them is five to twenty years, and he says the current approaches at other companies will stall without that, while still generating what he calls stupendous revenue. We have no way to check either forecast. What we can check is whether the generalisation gap he describes shows up in experiments a small lab can afford, and that is the experiment we would want to run first. If it is the research question of the next period, it is one that can be worked on from here.

Sources

  1. Dwarkesh Patel, Ilya Sutskever: We are moving from the age of scaling to the age of research (25 November 2025)