Ilya on next-token prediction: notes on a podcast that aged well
Sutskever told Dwarkesh Patel that predicting the next token well enough requires understanding the reality that produced it, and that this is why the objective need not cap out at human level. Reading notes on the argument and the objections we keep hearing to it.
The claim
Dwarkesh Patel put the standard objection to Ilya Sutskever yesterday in almost the words we would have used. If a model is trained to predict text written by humans, surely the best it can do is match the humans who wrote the text. Sutskever's reply was blunt. He said he challenges the claim that next-token prediction cannot surpass human performance, and then gave the reason, which is the part worth writing down.
His argument runs as follows. Predicting the next token well means understanding the underlying reality that led to the creation of that token. He was explicit that this is not a matter of statistics in the shallow sense. To compress the statistics of text you have to understand what it is about the world that produces those statistics. The text is a projection of processes, people, physics, arguments and mistakes, and the loss rewards a model that recovers enough of the process to anticipate the projection.
The thought experiment
The part of the interview that will get quoted is a thought experiment. If the base network is smart enough, Sutskever said, you can ask it what a person with great insight, wisdom and capability would do. Such a person may not exist in the training data, but a model that has learned the generative process behind human writing has a reasonable chance of extrapolating how that person would behave. The training set is a sample of what humans have done, and a good model of the sampler can be conditioned on regions the sample never visited.
We find this persuasive as an argument about the ceiling of the objective and unpersuasive as a statement about any current model. The gap between the two is the whole question. A model can have learned a rough process model that is good enough for typical text and still be nowhere near able to simulate a wiser version of its authors. Sutskever did not claim otherwise. He said the extrapolation is possible if the network is smart enough, which is a conditional, and we think the interview is best read as a bet on which direction the conditional resolves.
The objections at the time
The first objection we hear is that a loss on human text cannot reward anything humans did not write. The response in the interview is that the loss rewards modelling the process, and the process contains more than its outputs. A model that understands why a mathematician wrote a proof one way can in principle write a proof the mathematician did not. Whether the pretraining signal is strong enough to force that level of modelling, rather than a cheaper surface approximation, is the empirical question and the interview does not settle it.
The second objection is about reliability, and Sutskever raised it himself. He said that the main constraint on real-world economic value is reliability, and that models are already better at multi-step reasoning when allowed to think out loud. He was also candid that current understanding of what the models are doing is rudimentary. So the strong claim about the objective sits next to a modest claim about the present, and the two are consistent. A process that can in principle exceed its training distribution can still be too unreliable to use today.
The third objection is that improvement will come from somewhere other than the next-token loss, most obviously from reinforcement learning. Here Sutskever's comment was that most of the reinforcement signal already comes from AI, with humans used to train the reward function. That does not contradict the next-token thesis. The reward model is itself a product of pretraining, and the pretraining is what gives it anything to judge with.
Why we expect this to age well
The reason we think this interview will be cited in a few years is that it commits to a falsifiable position. If next-token prediction is an imitation ceiling, then models trained on it should plateau at roughly expert human level on tasks with plentiful human text, and no amount of scale or conditioning should get them past that. If Sutskever is right, we should see models produce correct outputs on problems where no human text contained the answer, and we should see those outputs improve with scale of the base model before any task-specific training.
The think-out-loud observation is a small hint in his favour. A model that reasons better when given more tokens to work with is a model whose predictions depend on a computation it is performing, not on a lookup. That is exactly what you would expect if the loss had pushed it towards a process model. It is not proof, and we would want to see the effect on problems with verified novelty before making more of it.
What we would want someone to try is the clean version of the extrapolation test. Take a base model with no instruction tuning, prompt it as Sutskever described, and score it on a set of problems constructed after the training cutoff with known solutions. Track the score against base model scale. If the curve rises, the ceiling argument is in trouble, and this interview will look like the moment the claim was made plainly enough to check.
Sources
From the foundation