The decade of agents: Karpathy on Dwarkesh, and nanochat
In one October week Andrej Karpathy told Dwarkesh Patel that agents will take a decade rather than a year and released nanochat, a full ChatGPT-style pipeline in about 8,000 lines that trains for around 100 dollars. Notes on why the two belong together.
Two releases in one week
On October 17 Dwarkesh Patel published a long interview with Andrej Karpathy, and the line that travelled was that this is the decade of agents rather than the year of agents. A few days earlier Karpathy had put nanochat on GitHub, a single repository that covers tokenisation, pretraining, fine-tuning, evaluation and inference for a small chat model, with a speedrun script that trains one on an 8xH100 node in a few hours. He calls it the best ChatGPT that 100 dollars can buy.
We have seen the two treated as separate news items. We think they are the same position stated twice. If you believe capable agents are ten years out, the most valuable thing you can do in 2025 is make sure a lot of people understand the whole stack well enough to do the intervening work, and a 100 dollar end-to-end pipeline is how you do that.
What the decade claim rests on
Karpathy is explicit that his timeline comes from watching predictions over about fifteen years in the field and noticing how consistently they run early. He is not claiming the problems are unsolvable. His phrasing is that they are tractable but still difficult. He lists what today's agents are missing, and the list is not short. They are not smart enough, they lack enough multimodal capability, they cannot use a computer well, they do not learn continually, and they cannot retain what they are told.
The continual learning point is the one he spends most time on and the one we found most persuasive. Everything a model knows was fixed at training time, and a session's worth of context is discarded at the end. He connects this to a problem he calls collapse. Ask a model for a joke and you get one of about three jokes. The distribution it produces is far narrower than the distribution it was trained on, and training a model on its own outputs compounds the narrowing. That closes off the easy route to continual learning through self-generated data, because the data is not diverse enough to learn from.
His critique of reinforcement learning is in the same spirit. He calls it terrible next to imitation learning, and the image he uses is sucking supervision through a straw. A model makes a hundred attempts at a maths problem, a few succeed, and every token in the successful trajectories is upweighted equally, including the tokens that contributed nothing. Humans, he notes, mostly do not learn intellectual skills this way. He still expects agents to arrive. His point is that the methods we have are inefficient enough that arriving will take years of work on the methods.
Why nanochat is the same argument
The repository is about 8,000 lines and its stated purpose is to be the full pipeline in one hackable place. The speedrun configuration trains on an 8xH100 node, which at around 3 dollars per GPU hour comes to roughly 24 dollars an hour, so a run of a couple of hours lands near the 48 dollar mark, and the README mentions a spot instance run at about 15 dollars. The reference point the README offers is GPT-2 in 2019, which it puts at around 43,000 dollars and 168 hours. The leaderboard target is the DCLM CORE score of the original GPT-2, 0.2565, and the record as we write is 0.2626 in 1.65 hours.
Every one of the gaps Karpathy listed in the interview is a training method problem, and training method problems get solved by people who can run the whole loop and change any part of it. A researcher who can only fine-tune a released checkpoint cannot work on continual learning, because the pretraining is where the representation was fixed. A researcher with nanochat can change the tokenizer, the data mix, the optimiser, the post-training recipe and the evaluation and see the result the same afternoon for the price of a dinner.
There is also a telling detail in the interview about how the code was written. Karpathy says the coding assistants were not net useful for nanochat because the architecture was unfamiliar to them. The models kept misunderstanding his gradient synchronisation, which he had written without the standard distributed data parallel wrapper, and tried to correct it toward the common pattern. His conclusion was that autocomplete is still the sweet spot. That is the decade claim in miniature. The agent could not do the novel part, and the novel part is where the field's progress lives.
The smaller predictions
A few other claims from the interview deserve a note in the ledger. He guesses that the cognitive core of a model, the reasoning capability stripped of memorised facts, could fit in about a billion parameters, and that today's trillion-parameter models are mostly carrying the cost of compressing low quality internet text. He describes the training corpus as largely garbage and expects that better curated data would let far smaller models reach the same capability.
On the economy he does not expect a discontinuity. He says AI is not identifiable in GDP data any more than computers or smartphones were, and that the growth curve is the same exponential the last few centuries of automation have produced. He restricts his own definition of AGI to knowledge work, which he estimates at 10 to 20 percent of the economy, and uses radiologists as the example of a profession that was predicted to vanish and instead grew. The framing he offers instead is an autonomy slider that moves gradually, task by task, with coding first because code is text-native and already has diffs and terminals built around it.
What we take from the pairing
The two releases give a coherent recommendation to anyone deciding what to work on. If the missing pieces are continual learning, sample efficient training, and diversity in what models generate, those are problems that live below the API, and they are cheap enough to touch now that a full run costs less than a conference registration. The intervening decade will be spent by people who can run the loop.
The thing we want to see is a nanochat-scale experiment on the collapse problem. Train a small model, sample from it at length, measure the diversity of what comes out against the diversity of the training set, and then try the obvious interventions. A result at that scale would say little about frontier models directly. It would tell us whether the narrowing Karpathy describes is a property of the method or of the size, and that is exactly the kind of question a 100 dollar pipeline exists to answer.
Sources
From the foundation