Open problems in mechanistic interpretability: thirty researchers write down what they do not know
Sharkey and 28 co-authors from 24 institutions sorted the gaps in mechanistic interpretability into methods, applications and socio-technical problems. Reading notes on the list, and on which items we would bet close first.
Who wrote it and why it is unusual
The review posted on 27 January has 29 authors, led by Lee Sharkey with Bilal Chughtai, Joshua Batson, Jack Lindsey and Jeff Wu among the first names, from 24 institutions including Apollo, Anthropic, EleutherAI, Google DeepMind, Goodfire, METR, MIT, Harvard, Tel Aviv, Melbourne and Timaeus. Getting that many people from competing labs to sign the same list of what the field cannot yet do is itself a result. Survey papers usually summarise what has worked. This one is structured as a catalogue of failures, with three top-level bins: methods, applications, and socio-technical problems.
We read it as a yardstick. A list this specific, dated this precisely, lets us check in a year or two which items moved. So these notes are partly a summary and partly a set of predictions we are willing to be wrong about.
Methods: decomposition is the load-bearing problem
The methods section is dominated by sparse dictionary learning, meaning SAEs and their relatives, because that is where the field's effort has gone. The authors list eight problems with it, and the first is the one that keeps coming up in our own work: reconstruction error. Their cited example is that inserting a 16-million-latent dictionary into GPT-4 raises the language modelling loss to what a model trained with only 10 percent of GPT-4's pretraining compute would achieve. Whatever the dictionary is describing, it is missing a lot of the model.
The rest of the list is a good summary of the field's private worries said in public. Dictionaries assume a linear representation in a nonlinear network. Sparsity is a proxy for interpretability and the two diverge, producing feature splitting and feature absorption. The features are treated as isolated directions with no account of their geometry. Dictionaries describe activations, and activations are not mechanisms, so they say nothing about the weights that compute them. And which concepts appear depends on the training distribution, so the concept you need for a safety question may never show up as a latent.
There is a separate set of problems about description and validation that we think deserve more attention than they get. Highly activating examples project human priors onto whatever direction you hand them. Attribution methods can be adversarially manipulated to produce any map. And the field's habit of evaluating on cherry-picked tasks means we rarely know whether a method works outside the paper that introduced it. The authors call for model organisms with known ground truth, multiple seeds and open weights, which is a request we would sign.
Applications: the gap between having features and using them
The applications section is shorter and reads as more tentative, which is honest. For monitoring, there is no framework yet for identifying unsafe cognition mechanistically, and scaling to frontier models is untested. For control, SAE latents do not always contain the concept you want to edit, and unlearning results suggest that latents and human concepts line up worse than we would like. For predicting behaviour in new situations, mechanistic methods have not yet beaten simple baselines. And for microscope AI, the dream of reading knowledge out of a model, the same data-dependence problem returns: you find what the training distribution put there.
The cross-cutting complaint is about competition. Interpretability methods are rarely tested head-to-head against non-interpretability baselines on an engineering goal. If a probe or a fine-tune achieves the same monitoring accuracy at a tenth of the cost, that should be in the paper. It usually is not.
Socio-technical: what counts as understanding
The third bin is the one most reviews skip. It asks what interpretability results are for in governance, and notes that no clear path yet runs from a circuit paper to a regulatory requirement. It asks what understanding even means for a system this large and whether mechanistic decompositions will ever align with the concepts humans use. And it raises the evaluation standards problem: the field has no agreed level of rigour, and sanity-check failures of the kind Adebayo and colleagues found in saliency maps in 2018 could easily recur for current methods.
This section could have been longer. The question of who decides which concepts get labelled, and how much power that gives the labeller, is raised in a paragraph and then dropped. We would like a follow-up on just that.
Which of these we expect to close
Here is our scorecard for checking back on. The decomposition problems about architecture, meaning features that span layers and attention heads, look tractable because people are already training dictionaries across layers. The reconstruction error problem looks harder, because the fix might be that dictionaries are the wrong object rather than a better dictionary. The model organism and benchmark requests should be met within a year because they require organisation rather than a new idea. The competitive-baseline problem will close only if reviewers start demanding it.
The socio-technical items we do not expect to close, because they are not the kind of problem that closes. What we would hope for is that when the next version of this list is written, the methods section is shorter, the applications section has numbers in it, and the socio-technical section is longer. If the reverse happens, that is also worth knowing.
Sources
From the foundation