Why we took notes on this one

Dario Amodei sat down with Dwarkesh Patel last week for a long interview, and it is the most candid public account we have heard from a frontier lab head about what they do and do not understand. We are writing this partly as a reading guide and partly as a ledger. Several of the claims have dates attached, and the useful thing to do with a dated claim is to write it down now, with the timestamp, and check it later.

The interview was recorded in early August 2023 and runs a little over two hours. The timestamps below are approximate and come from the published transcript.

The admission about scaling

The opening stretch, around the first three minutes, is where Amodei says the thing that gives this post its title. Asked why adding parameters and data produces such smooth improvement, he answers that we still do not know. He offers the power law correlations that appear everywhere in the scaling literature and a loose analogy to the fractal dimension of a manifold as gestures toward an explanation, and then says explicitly that these are gestures and that the field lacks a theory.

He then draws a distinction we think is the most important one in the interview. Loss is predictable, in his words to several significant figures, as a function of compute. Specific abilities are not. A capability that is absent at one scale can appear abruptly at the next, and nothing in the loss curve tells you in advance which capability or which scale. The smooth curve and the jumpy capabilities are both true at once, and the second is what makes deployment and safety hard even when the first makes training plannable.

The line that follows, that the models just want to learn and you get the obstacles out of their way, is the one that will be quoted. We read it as a description of what the last five years of engineering have felt like from the inside rather than as a theory. Remove the bottleneck, whether that is data, precision, or a bad architecture, and the loss goes down again. Nobody has explained why the ceiling keeps not arriving.

The predictions, dated

Around the 28 minute mark Amodei gives a timeline for models that behave like a generally well educated human across most tasks. His figure is two to three years, so somewhere in 2025 or 2026, with a caveat that safety thresholds might delay when such a model is released rather than when it could be built. He is careful to separate the technical capability from the product.

At the 38 minute mark he gives the same two to three year window for a different and much darker claim, that models may be able to provide meaningful uplift for large scale biological attacks. He says the current models cannot fill the key missing pieces in such a workflow and that the trend line is what worries him. Later, discussing the economics, he says spending on the largest training runs is likely to grow by something like a factor of 100, and puts the cost of GPT-4 and Claude 2 at roughly 100 million dollars each. Those two figures together imply runs in the billions within a few years.

There is also a rough description of the gap between capability and deployment. He describes Claude today as roughly an intern in most areas, with a few specific areas where it is better, and spends several minutes on the economic friction between what a model can do in a demo and what it changes at scale. That is the prediction we find hardest to score, because friction is measured in years and in ways nobody has agreed on.

Data, brains and interpretability

Two smaller threads are worth pulling out. On data, Amodei is unbothered. Asked whether the supply of text is the binding constraint, he says there are many sources of data in the world and many ways to generate it. He does not elaborate on what synthetic data would look like at frontier scale, which is the part we would have pressed on.

On sample efficiency he concedes the odd shape of the current models. By his rough numbers they are two to three orders of magnitude smaller than a human brain and are trained on three to four more orders of magnitude of data. A person has seen on the order of hundreds of millions of words by 18. A model has seen hundreds of billions to trillions. Something about how the two systems learn is different, and the interview does not pretend to know what.

The interpretability section, from roughly 47 to 52 minutes, frames mechanistic interpretability as an X-ray rather than a surgical tool. The goal is a way to inspect what a model is doing that is independent of behavioural evaluation, so that an evaluation which passes and an internal inspection which passes are two different pieces of evidence. He does not claim it solves alignment on its own, and he says the lab is currently bad at controlling the models despite its alignment work, which is a striking thing to say on the record.

What we will check

The scorable claims are these. Models that look like a well educated generalist by 2025 or 2026. A hundredfold increase in spend on the largest runs. Meaningful biological uplift within the same window. Interpretability producing inspection tools that add evidence beyond evaluations. We are putting a note in the calendar for the end of 2024 to come back to this list with what has actually shipped, and we will be most interested in whether the capability jumps he describes as unpredictable turn out to have been predictable in hindsight by anyone at all.

The claim we most want someone to work on is the one he cannot answer. A theory of why loss falls as a power law in compute, even a partial one that explained a single regime, would tell us more about where this ends than any timeline.

Sources

  1. Dwarkesh Patel, Dario Amodei interview (August 8, 2023)