Sutton says LLMs are not bitter-lesson-pilled
The author of the Bitter Lesson went on a podcast and said the field that quotes him constantly is doing the opposite of what the essay recommends. He is partly right, and the part he is right about is the part nobody wants to hear.
What Sutton said
Richard Sutton wrote the Bitter Lesson, the 2019 essay that most of the field now cites as the justification for scaling. This month he sat down with Dwarkesh Patel and said that large language models are not an example of it. "Large language models are learning from training data," he said. "It's not learning from experience." His view is that they "are about mimicking people, doing what people say you should do. They're not about figuring out what to do."
The argument has three parts. First, an LLM has no goal in his sense. "If there's no goal, then there's one thing to say, another thing to say. There's no right thing to say." Second, it does not have a world model. "They have the ability to predict what a person would say. They don't have the ability to predict what will happen." Third, it cannot learn continually, which he defines plainly. "If you need to learn continually, continually means learning during the normal interaction with the world." He expects "systems that can learn from experience" to be "much more scalable" and to supersede the current approach, which would make LLMs the human-knowledge-laden method that the bitter lesson says eventually loses.
The reaction has been loud, and mostly hostile. We want to work out which parts of the hostility are justified, because we think the field is having two arguments at once and confusing them.
The argument he loses
Zvi Mowshowitz wrote the most thorough response, and on the definitional points we think he wins. His charge is essentialism. Sutton declares that LLMs cannot have goals or world models because they lack some property, and then when they behave as if they had both, the behaviour is dismissed as not counting. Zvi's example on goals is that "next token prediction does influence the world in training, because feedback will change the model's weights." On the claim that infants do not learn by imitation, he points out that newborns imitate facial expressions within hours of birth, which is not a controversial finding.
Zvi's sharpest diagnosis is methodological. Sutton "isn't applying the theory to anything." Rather, "he is inside the theory and interpreting everything according his understanding of it." That rings true. When you have spent forty years building a formalism where intelligence is defined as reward maximisation through action, a system that reaches useful behaviour some other way looks by construction like not-intelligence. The definitional argument is not one Sutton can win, because he is not arguing from evidence about the models. He is arguing from the axioms of his field.
The argument he wins
And yet. Strip out the words "goal" and "intelligence" and "world model," each of which Zvi is right to call a word game, and there is a plain empirical claim left over. Current LLMs do not learn during deployment. Their weights are frozen at release. Everything they know, they knew when training stopped, and everything they learn in a conversation is gone when the context window closes. Zvi concedes this, and calls it a genuine architectural difference. So does everyone who has tried to build a system that gets better at a user's specific job over months.
That is the part of Sutton's critique that is uncomfortable, because the field has a habit of quoting the Bitter Lesson while doing something the essay explicitly warns against. The essay says that methods which build in human knowledge win in the short run and lose to methods which learn from computation in the long run. Pretraining on the entire written output of humanity is the largest injection of human knowledge in the history of the discipline. It is very effective. It is also, by the essay's own logic, the thing that gets superseded.
The field's reply, which Zvi makes, is that scaling learning has beaten hand-engineering, and pretraining is scaled learning, so the Bitter Lesson is honoured. We think that is true and beside Sutton's point. His point is about where the learning signal comes from. In the settings the essay was written about, search and self-play, the signal came from the environment. In pretraining, it comes from people. Whether that distinction matters for the long run is exactly the open question, and we do not think either side has evidence that settles it.
What the debate is actually about
The two arguments are these. The first is whether current systems deserve the words we use for them. That one is unwinnable and mostly uninteresting, and both Sutton and his critics spend too much time on it. The second is whether a system that learns only from human-generated data can keep improving once it has consumed most of it, or whether continued progress requires learning from its own interaction with the world. That one is empirical, it is the question Sutton is actually raising, and it has a plausible answer in either direction.
For what it is worth, the reasoning models released this year lean Sutton's way. The reinforcement learning stage that gives them their reasoning ability is learning from a verifiable signal, from whether the answer was right, rather than from what a person said. That is closer to experience than to imitation, and it is where the recent capability gains came from. If that trend holds, the field will have quietly agreed with Sutton while insisting it did not.
What we would want someone to try
The experiment we want is a controlled comparison of continual learning against context. Take a fixed model, a long-running task with feedback, and measure whether updating the weights during deployment beats the same model with an ever-growing context and retrieval. The field has assumed the second is good enough because it is easier to ship. Sutton's claim is that it is not, and that the gap will show up as a ceiling. That is a testable claim, and a small group can test it, and we would rather read that result than another round of arguing about what a goal is.
Sources
From the foundation