How LLMs actually think: notes on the Sholto and Trenton conversation
A long Dwarkesh Patel episode with Sholto Douglas of Google and Trenton Bricken of Anthropic has become the thing people send newcomers. Reading notes on what it claims about in-context learning, long context, superposition and the daily work of research.
Why this one gets passed around
The episode went out on March 28 and within days three different people had sent it to us with some version of 'this is the one to listen to'. Sholto Douglas works on Gemini at Google. Trenton Bricken works on interpretability at Anthropic and is one of the authors of Towards Monosemanticity. It runs for hours, and the reason it travels is that both of them talk about the mechanics of their work at a level between a paper and a tweet, which is the level where a person trying to enter the field learns the most.
These are our notes, sorted by claim rather than by timestamp, with our own view on how much weight each claim can bear.
In-context learning as something like gradient descent
Douglas describes in-context learning as behaving like gradient descent carried out through attention, and points to work showing a correspondence between n layers of in-context processing and n steps of gradient descent on the examples in the prompt. That is a specific mechanistic claim, and it is one that has been demonstrated for small models on constructed tasks. Whether it is the right description of what a frontier model does when it picks up a pattern from your prompt is open, and we would treat it as a useful picture rather than an established mechanism.
The more surprising part of his position is about long context. He argues that context windows in the million-token range are underrated, because a model with that much relevant material in front of it gets dramatically better at next-token prediction in ways that previously required a bigger model. The evaluation he cites is learning an obscure human language from material placed in context that the model had not seen in training. If that holds up generally, then a lot of what we have been buying with parameters can be bought instead with retrieval and context, and the economics of that trade are very different.
Superposition, in the words of someone who works on it
Bricken's explanation of superposition is the clearest short account we have heard. A model represents far more features than it has dimensions by letting features overlap in a high-dimensional space, on the bet that most features are rarely active together, so the interference is tolerable. The consequence is that individual neurons are polysemantic and hard to read. The Towards Monosemanticity approach is to project activations into a much larger space with a sparsity penalty, and the claim is that when you do this the features come out clean.
What we found useful was his framing of intelligence itself as hierarchical association, pattern matching stacked on pattern matching, and his willingness to say that this might be all there is. It is a substantive position and it has consequences for what interpretability can hope to find. If cognition is associations at many levels, then a dictionary of features plus a map of how they combine is a complete description, and the project is finite. If there is something else, the dictionary is a start and no more.
What a good researcher does all day
The part we would make required listening is the section on the shape of the work. Douglas describes research as a cycle: have an idea, prove it out at a few points along the scale, and work out what went wrong. He is direct that the bottlenecks are compute, taste in choosing what to try, and the difficulty of reading noisy scaling-law data, rather than the engineering itself. Both guests come back to the idea that the limiting resource is judgment about what experiment is worth the GPUs.
There is a quantitative aside we keep thinking about. Asked what more compute would buy a research group, the estimate offered is that ten times the compute yields roughly five times the research progress. That is a guess by two people, and we would not put it in a paper. But it is a guess by two people who allocate compute for a living, and the sublinear shape is the part that matters. Ideas and people run out before chips do.
They also talk about agents. Their claim is that useful agents arrive when reliability crosses into the nines, and that the near-term version will be several specialised models talking to each other in text rather than one model trained end to end for long tasks. That is a testable prediction with a date attached to it, roughly, and we will know within a year or two.
What to be careful about
Some of the numbers thrown around are rough and should be treated that way. The comparison between synapse counts in the brain and parameter counts in a model is illustrative and nothing more. The training cost estimate for GPT-4 is a public guess. The correspondence between in-context learning and gradient descent is a result on toy models being applied by analogy to large ones. None of that is a criticism of the conversation, which flags its own uncertainty, but the numbers will be repeated without the flags.
The other thing to hold onto is that both guests work at labs with a view. Douglas's enthusiasm for long context and Bricken's confidence in dictionary learning are each the position of the team they are on. That does not make them wrong. It means that if you are using the episode to decide what to work on, you are hearing two answers out of several.
What we would do with it
If we were starting in the field we would listen to it twice, once for the ideas and once for the descriptions of daily practice, and then pick one claim to try to falsify at a scale we can afford. The long-context-versus-parameters claim is the one we would choose, because it is the one with the largest consequences if true and it can be tested on models we can run.
Sources
From the foundation