The data black hole: Dwarkesh on sample efficiency and learning on the job
Two essays from Dwarkesh Patel this month argue that a million-fold gap in sample efficiency is the unsolved problem, and that on-the-job learning from deployment data is the next big breakthrough. Reading notes, a check against his own essay from a year ago, and what a small lab could actually test.
The argument in the first essay
Dwarkesh Patel published two linked essays this month. The first, on June 19, makes a claim about where progress comes from. His position is that data quantity and quality drive capability, that architecture and hyperparameters matter much less than they appear to, and that this is why open models trail closed ones by only about four months. The evidence for that lag is not laid out in detail, so we are treating it as his estimate rather than a measurement.
The number he builds the piece around is the sample efficiency gap. A human encounters on the order of 200 million tokens of language between birth and adulthood. Frontier models train on tens to hundreds of trillions. That is roughly a million-fold difference, and he pairs it with a familiar contrast in driving, where a teenager needs around 20 hours and autonomous vehicles needed millions of hours of demonstrations. His definition of intelligence for the purpose of the essay is how much data you need to see in a domain before you operate fluently in it.
The part we found most useful was the description of what the data pipeline now looks like. He cites job listings from Mercor and Surge for document specialists, lawyers writing realistic M&A diligence, and consultants building market research templates. The claim is that every skill the models acquire is being bought as hundreds of expert trajectories per skill, and that this grafted-together structure cannot reach general competence by scaling alone. He also notes that if you trained compute-optimally and wanted to cut data needs, even infinite parameters would only reduce them by about a factor of ten.
The second essay: learning on the job
The June 26 piece proposes the answer. Its subtitle is that labs are throwing away the most valuable data, and the data in question is deployment. He puts 30 to 50 percent of a lab's compute in inference and says that compute is currently not doing anything to improve the model. The information that only appears in use, such as how a specific organisation works, what a specific user wants, and where the model fails in practice, is lost at the end of every session.
The mechanisms he offers are concrete enough to argue with. One is on-policy self-distillation, where the base model is trained to match the predictions of a copy of itself that has accumulated a session's worth of context, so that in-context learning is converted to weight changes without needing a verifiable reward. Another he calls dreaming, in which the model builds simulations to rehearse a skill before doing it. A third is architectural, with longer contexts and KV cache compression so that a session's learning can be retained. He points to Cursor's Tab model learning online from more than 400 million requests a day as an existing example, and contrasts the roughly 320 KB a KV cache grows per token with the 0.075 bits per token a model absorbs in training, a gap he puts at 35 million fold.
Checking it against the year-old version
Patel has made this argument before. On June 2 last year he wrote that continual learning was the reason he did not think AGI was near. The illustration was teaching a child the saxophone by having one student try, writing detailed notes on the failures, and handing the notes to the next student. His bets were 50 percent that an AI could do a small business's taxes end to end by 2028 and 50 percent that it could learn on the job as well as a human across white-collar work by 2032.
A year on, the diagnosis is the same and the proposed cure has gotten more specific. What has not changed is the absence of a demonstration. The new essay mentions no result showing a deployed model that improves through use on a task and keeps the improvement without regressing elsewhere. Cursor's Tab model is the closest, and it is a narrow completion model with a dense, free reward signal, which is the easy case. The bits per sample argument he references from his own earlier writing, that RL absorbs less information per sample than supervised learning and may therefore forget less, is an interesting claim about a mechanism and would also need a measurement.
So the honest summary is that the essays describe the problem well, describe a research direction that several labs are presumably already pursuing, and provide no evidence that the direction works. That is fine for an essay. It should not be mistaken for a forecast with support behind it.
What a small lab could test
The self-distillation proposal is testable at our scale, and that is the reason we are writing this up. Take an open model of a few billion parameters. Give it a task family where a session of interaction produces context that helps, such as a codebase with local conventions or a customer with stated preferences. Run the veteran copy with the context, train the base copy to match its distribution on the next session's prompts, and measure three things. Does the base model reach the veteran's accuracy without the context. Does it lose accuracy on a held-out benchmark. And how many sessions of this before the second number moves.
That experiment would put a number on the thing the essays leave open, which is the exchange rate between on-the-job learning and forgetting. If it is favourable for a small model on a narrow domain, the argument that labs are wasting their inference compute gets much stronger. If it is unfavourable even there, then sample efficiency is still the black hole and the deployment data was never the way out.
Sources
From the foundation