Reading 'Welcome to the Era of Experience'
Silver and Sutton argue that models trained on human data are approaching a ceiling and that the next generation of agents will learn predominantly from their own experience. Reading notes on the four properties they propose, set against the RL-trained reasoning models shipping at the same time.
The claim
David Silver and Richard Sutton have posted a short position paper, a preprint of a chapter for an MIT Press book called Designing an Intelligence, arguing that the field is at the threshold of what they call the era of experience. The argument in one paragraph: training on human-generated data has produced broad competence, but in mathematics, coding and science the useful human data is close to exhausted, the pace of progress from supervised learning on it is slowing, and the insights that matter most lie beyond current human understanding and so cannot be in the data at all. The way forward is agents that generate their own data by acting in an environment and learning from what happens.
They sketch three eras. An era of simulation, running from Atari through AlphaGo and AlphaZero, in which RL agents mastered closed problems with clean rewards but never crossed into open-ended ones. An era of human data, from GPT-3 through ChatGPT, in which the field largely discarded experiential RL in favour of imitation and preference tuning. And an era of experience, which they say has already begun with AlphaProof and DeepSeek-R1, and which will eventually produce experiential data that dwarfs the scale of human data used in today's systems.
The four properties
The substantive part of the paper is a list of four ways experiential agents differ from current systems. Streams: agents live in a continuing stream of experience over months or years rather than short episodes, carrying information across the stream and pursuing goals set far into the future. Actions and observations: agents act in the world through the same interfaces humans use, APIs, browsers, sensors, instruments, rather than only exchanging text with a user. Grounded rewards: the signal comes from consequences in the environment, heart rate, exam results, carbon dioxide levels, tensile strength, rather than from a human judging the action before its effects are known. And non-human reasoning: agents discover their own ways of thinking, possibly in non-linguistic representations, rather than imitating human chains of thought.
The third of these is where the paper does its most careful work. The authors acknowledge that optimising a single grounded signal does not obviously give you a system that can be steered toward arbitrary user goals. Their proposal is a learned reward function that takes the user's stated goal and the agent's interactions with the environment as input, combines grounded signals accordingly, and is itself adapted by user feedback in a bi-level optimisation. A small amount of human data steering a large amount of autonomous learning. That is a design, and it can be argued with, which is more than most position papers offer.
Prophecy or nostalgia
The paper reads differently depending on which shipping model you hold it against. Against a chat assistant tuned on human preferences, it is a prophecy. Against DeepSeek-R1, released in January and cited by the authors as evidence that the transition has begun, it reads as a description of something that already happened. R1's reasoning was trained with reinforcement learning against verifiable outcomes, and the authors quote DeepSeek's own line that rather than teaching the model how to solve a problem, they provide the right incentives and it develops strategies on its own. On the reward axis, the era of experience is the current era.
On the other three axes the case is weaker. The reasoning models shipping this spring still think in human language, the paper itself notes this and argues it is unlikely to be optimal, but there is no shipped system that reasons in anything else. Their episodes are still short. And their actions are mostly text, with computer-use agents cited as early exceptions. So the honest scorecard is one property largely arrived, three still aspirational, and the paper's own figure, which shows the field's attention to RL dipping through the ChatGPT years and climbing again, is a claim about attention rather than about results.
There is also a strand of the paper that reads as an argument the authors have been making for a long time. The section on RL methods says the human-data era threw out the baby with the bathwater, that RLHF sidestepped value functions, that strong priors from human data lessened the need for exploration, world models and temporal abstraction. That is a reasonable reading of history. It is also exactly what two of the people who built those methods would say, and the paper does not engage with the possibility that the human-data detour was what made general-purpose agents possible in the first place.
The safety section is better than we expected
Position papers of this kind usually treat safety as a paragraph at the end. This one spends a page on it and makes three arguments worth recording. First, an agent that learns from its environment can notice when its behaviour is causing human concern and adapt, where a fixed system cannot. Second, a reward function that is itself adapted through experience can have misalignments corrected incrementally by trial and error, the paperclip example being modified before all the resources are consumed. Third, and this is the one we find most persuasive, learning from physical experience is bounded by the time physical experiments take, and that acts as a natural brake on the pace of self-improvement. A drug still needs a trial.
Each of those cuts both ways. An agent that adapts to human concern is also one whose goals shift in ways that are hard to inspect. Correcting a reward by trial and error assumes the errors are survivable. And the physical brake applies only to agents whose experience is physical, which is not where most of the current effort is going. The authors say plainly that moving away from human data and human modes of thinking may make systems harder to interpret. We would have liked a sentence on what they propose to do about that.
What we would try
The paper is a research agenda more than a forecast, and the testable version of it is small. Take a reasoning model trained on verifiable rewards, give it a long-horizon environment with a grounded signal that is not a correctness check, something like an experimental loop where the reward arrives days later, and see whether the same RL recipe still works when the reward is slow and noisy rather than fast and binary. Our guess is that most of the machinery the authors want to bring back, value functions, exploration, temporal abstraction, becomes necessary exactly there, and that nobody has yet shown it working at language-model scale. That would be a useful thing to fail at in public.
Sources
From the foundation