From mechanistic to compositional interpretability: a category theorist's proposal
A new paper recasts mechanistic interpretability as an optimisation over faithfulness and description length, with existing methods as special cases. Reading notes on what the formalism buys and what it leaves for later.
The complaint the paper starts from
Mechanistic interpretability has no shared definition of what a good explanation is. Ward Gauderis, Thomas Dooms, Steven Homer, Kola Ayonrinde and Geraint Wiggins open their May 9 paper with that observation, and the sentence they build on is blunt: without a formal framework, mechanistic explanations cannot be objectively verified, compared, or composed. Anyone who has tried to decide whether two circuit papers on the same model are describing the same mechanism will recognise the problem.
Their answer is a category-theoretic framework they call compositional interpretability. We are not a category theorist and we will not pretend the paper was an easy read. What we can do is report what the machinery is for, which turns out to be more practical than the title suggests.
Two maps that have to agree
The central object is a pair of mappings. One is syntactic and describes how a model is wired together out of components. The other is semantic and describes what each component computes. The paper's requirement is that these two maps commute: if you decompose the model structurally and then interpret the pieces, you must land in the same place as if you interpret the whole model's behaviour directly. That commuting diagram is the consistency condition that ties a decomposition to observed behaviour.
Faithfulness, in their Definition 3.1, is then the ability to reconstruct the semantics of each explained component from the interpretation, within some tolerance measured by a metric on semantic elements. This is stricter than the usual behavioural notion, where an explanation only has to reproduce input-output behaviour. It has to be right at the component level, and the tolerance is what lets you abstract lossily without losing the right to call the explanation faithful.
Complexity gets split in two
The second half of the objective is complexity, and the paper's decomposition of it is the part we found most useful. Compositional description length has a representation term, the cost of describing the wiring plus the semantics of each component, and an interpretation term, the difficulty of lining those component semantics up with human concepts. The representation term depends only on the model. The interpretation term isolates the subjective part. Theorem 3.3 shows the sum upper-bounds the minimal description length of the explanation.
Interpretability then becomes constrained optimisation: minimise complexity subject to faithfulness within tolerance. The tool for doing that is what they call compressive refinement, a functor that restructures a model's syntax while preserving its semantics, and is compressive when it lowers representation complexity. Their worked example in Appendix D is small enough to check by hand. A linear classifier that costs 2,096 bits to describe drops to 1,744 bits once shared representations are disentangled by a refinement.
Proposition 4.2, the parsimony criterion, is the decision rule: a refinement produces a better explanation if and only if the drop in representation complexity outweighs the rise in interpretation complexity. That acknowledges something practitioners already know. You can compress a model into pieces that are simpler in bits and harder to name.
Existing methods as points on a dial
Table 1 in Appendix C is where the paper earns its subtitle. It lays out mechanistic methods along a spectrum defined by how much constraint is placed on the refinement. At one end, post hoc methods like saliency maps use no syntax at all. Architectural methods fix the atoms to attention heads and neurons, which is an identity refinement. Transcoders are piecewise, refining each component independently and therefore missing cross-layer patterns. Attribution graphs built on crosscoders only require the whole diagram to be faithful, so component correspondence can be dropped. Weight-sparse transformers sit at the far end, with no refinement constraint because the model itself is analysed directly.
That ordering explains a set of trade-offs that usually get argued about case by case. Stricter constraints on the refinement preserve traceability back to the original architecture and miss mechanisms that span components. Looser constraints expose more structure and risk explanations that no longer correspond to anything you can point at in the weights. The paper also positions itself relative to causal abstraction: that line of work assumes a high-level explanation is known and aligns it top-down, whereas compressive refinement searches bottom-up for decompositions that a causal alignment could then use.
Is formalism what the reproducibility problem needs?
The field's reproducibility problem is real. Two groups find different circuits for the same task, and there is no agreed way to say which is better or whether they are the same circuit seen twice. The paper's framework gives you a scoreboard: faithfulness as a measurable reconstruction error, complexity as a description length. If both groups reported those two numbers, you could at least compare them.
Our hesitation is about what the framework leaves as future work. Section 6 says plainly that automated compressive refinement requires faithfulness and complexity to be efficiently measurable and optimisable, and that the description and validation stages need separate treatment. The interpretation complexity term, the one that captures how hard components are to name, is exactly the part that has resisted measurement for years. A framework that defines the score is progress. A framework that also tells you how to compute the score at the scale of a frontier model is the thing we actually need, and this paper does not claim to be that.
What we would want to see next is a small, ugly test. Take two published circuit analyses of the same behaviour in the same open model, recast each as a refinement in this framework, and compute the two numbers. If the exercise is tractable and the numbers disagree with the field's informal ranking of the two analyses, that would be the most interesting result of the year. If it is not tractable, that tells us where the framework needs to go before it can referee anything.
Sources
From the foundation