'The Urgency of Interpretability': a lab CEO asks for a five-year head start
Dario Amodei's April essay argues that interpretability could be reliable by 2027 if the field moves now, and that transformative models might arrive before it does. Reading it as a research agenda and asking what a small independent lab can do about a race a frontier lab says it might lose.
The argument in one paragraph
Dario Amodei's essay makes a claim about timing. On one side, he expects models he describes as a country of geniuses in a datacenter to be possible around 2026 or 2027. On the other, he thinks interpretability could reliably detect most model problems by 2027 if the field accelerates, and that a mature version, the MRI for AI in his phrase, is a five to ten year project. The two curves cross uncomfortably close together, and the essay is a request for everyone to push the second one left.
We are going to read it as a research-agenda document rather than as a statement of corporate position, because it is more useful that way. Stripped of the framing, it says which problems are solved, which are half solved, and what a reliable diagnostic would need to do. That is a map, and maps are what an independent lab needs.
What he counts as progress
The essay lists four results as the evidence that the field has moved in the last year or so. Sparse autoencoders found over 30 million features in Claude 3 Sonnet, with the essay guessing the true count in even a small model could be a billion or more. Circuit tracing mapped groups of features that carry out steps of reasoning, with the Dallas to Texas to Austin example standing in for the general case. Golden Gate Claude showed that a single feature can be turned up to change behaviour. And an internal auditing game had a red team plant an alignment flaw in a model and blue teams find it, some of them using interpretability tools.
Notice what kind of evidence that is. Three of the four are demonstrations that the tools can see something. Only the auditing game is a demonstration that the tools can find a problem someone hid on purpose, and that is the task the whole essay is about. The gap between seeing features and reliably catching deception is where the five to ten years go.
What an MRI would have to do
The diagnostic Amodei wants would scan a model before deployment and flag deception, power-seeking tendencies, and jailbreak vulnerabilities, among other things. The list is worth reading as a specification. It implies a tool that works on a model it has not been trained on, that produces a verdict a non-specialist can act on, and that holds up when a model has been trained against the tool and is trying to hide from it.
None of the four results above has any of those three properties yet. Feature dictionaries are trained per model. Circuit tracing is done by hand by researchers who already know what they are looking for. The auditing game was played by people from the lab that built the model. Turning demonstrations into a scanner is an engineering programme, and the essay is honest that it does not exist.
The institutional asks
The recommendations are addressed outward. Anthropic says it is doubling its own interpretability investment with 2027 as the target for reliability. Amodei asks Google DeepMind and OpenAI to put more resources in. He wants neuroscientists and academics recruited, and the essay ends with job postings. For governments he suggests light transparency rules requiring labs to disclose their safety practices, and he argues that chip export controls to China buy a one to two year security buffer in which interpretability can catch up before parity.
The export control paragraph is the one place the document stops being a research agenda and becomes something else. We will only say that a safety argument for a trade policy should be evaluated as a trade policy, and that nothing about the technical content of the essay depends on it.
What an independent lab can contribute
Here is where we think a small lab fits. The big labs will build the scanners for their own models, and they will do it with access to weights, training data and training runs that no one else has. What they will not do well is validate each other's tools. A diagnostic that only its author has run is not yet a diagnostic. Independent replication of feature and circuit results on open-weight models is cheap, and it is the step that converts a lab's internal demonstration into a method the field can trust.
The auditing game is the other thing we can do. The version in the essay was played inside one company. A public version, where one team hides a flaw in an open model and anyone can try to find it with tools of their choosing, would tell us what the tools can catch when the finder is not also the builder. That is closer to how a pre-deployment scan would be used, and it produces a leaderboard where the score is a number that matters.
Finally, the timeline in the essay is a forecast, and forecasts can be tracked. If reliable detection is supposed to arrive by 2027, someone outside the lab should write down now what reliable would mean in measurable terms and check in each year. We would rather we be the ones holding the clipboard than have the deadline pass unmeasured.
Sources
From the foundation