Tracr and the case for ground-truth transformers
DeepMind's Tracr compiles programs written in RASP into the weights of a standard decoder-only transformer, so that the circuit inside is known before anyone looks. Why a laboratory with answer keys matters for interpretability, and what compiled models still fail to capture about trained ones.
The problem with grading your own homework
Every mechanistic interpretability result we have read this year has the same structural weakness. Someone finds a circuit in a trained model, describes what it computes, and offers evidence from ablations or activation patching that the description is right. The evidence is usually persuasive. It is never conclusive, because there is no ground truth. Nobody knows what the model actually computes, so nobody can say whether the description is complete, partial, or a convincing story about the wrong mechanism.
A paper posted to arXiv on January 12 by David Lindner, Janos Kramar, Sebastian Farquhar, Matthew Rahtz, Thomas McGrath and Vladimir Mikulik at DeepMind takes a direct route around this. Instead of training a transformer and then trying to read it, they write the algorithm down and compile it into transformer weights. The result is a model whose internal mechanism is known by construction, and which can therefore be used to test whether an interpretability method recovers what is actually there.
RASP, the language Tracr compiles
The source language is RASP, a small programming language designed by Weiss and colleagues to describe computations a transformer can express. It has three kinds of operation. Elementwise operations apply a function to every position of a sequence independently, which corresponds to what an MLP layer can do. Select operations build a boolean matrix over pairs of positions from a comparison predicate, which is an attention pattern. Aggregate operations use a selector to average values across the selected positions, which is what attention does with its value vectors.
With those three primitives you can write recognisable programs. The paper's examples include counting the frequency of each token in a sequence, sorting a sequence, and checking whether a string of parentheses is balanced. Each is a few lines of RASP. Each also has an obvious transformer implementation in the head of anyone who has thought about attention for long enough, and the point of Tracr is to make that implementation concrete rather than imagined.
How the compiler lays out the residual stream
Tracr turns a RASP program into a computation graph, then assigns each operation to a layer so that every operation runs after the ones it depends on. Elementwise operations become MLP blocks, with lookup tables for categorical inputs and piecewise approximations for numerical functions. Select and aggregate pairs become attention heads whose query and key weights implement the predicate and whose value and output weights carry the aggregated quantity.
The part that makes the models readable is the residual stream layout. Every variable in the program is given its own dedicated subspace of the residual stream, and values are stored either as one-hot categorical encodings or as a single numerical dimension. Nothing overlaps. If you want to know what a compiled model represents at layer three, you read the labels the compiler attached to the residual dimensions. That is the answer key, and it is exact.
Where compiled models diverge from trained ones
The authors are careful to say that a compiled model is a laboratory, and a laboratory is not the field. Three differences stand out. Trained models are not laid out with one variable per subspace. They pack far more features than they have dimensions, in superposition, and reading a trained residual stream means untangling that. A compiled model has no superposition unless you add it, and so an interpretability method that works on Tracr output has only been shown to work in the easy case.
The paper addresses this directly with an experiment. They take a compiled model and compress its residual stream with gradient descent, training a smaller model to match the behaviour of the compiled one in fewer dimensions. The optimiser discovers overlapping representations on its own, which produces a model that has known structure at the program level and superposition at the representation level. That is a much better test object than the clean compiled model, though still an artificial one.
The other two gaps are about expressiveness. Tracr currently supports one attention head per select-aggregate operation, so it cannot produce the kind of multi-head cooperation that shows up in trained models. And its numerical MLPs are approximations with bounded but nonzero error, which means a compiled model computing a numerical function does so slightly wrong, while a trained model computing the same function might use an entirely different and messier decomposition.
What we would use it for
The obvious use is as a benchmark for circuit discovery. Compile a set of programs, run a discovery method on each compiled model, and score it on whether it recovers the known circuit. Several groups have methods that they believe find circuits in GPT-2. None of them has been graded against a model where the circuit is known. Tracr makes that grading possible, and we would expect the first round of results to be humbling.
The second use is subtler. Compiled models let you ask what a particular algorithm looks like when implemented in a transformer, which is a question you can then take to a trained model as a hypothesis. If you suspect a trained model is sorting, compile a sorter, look at its attention patterns and residual structure, and go looking for the same signature. Where the trained model differs, you have learned something about how training solves problems differently from a programmer.
What the paper does not do, and what someone should try, is train a model from scratch on the same tasks the compiled models implement and compare the two mechanisms side by side. The compiler is open source. The training runs are small. The comparison would tell us how far the laboratory is from the field, which is the number every result built on Tracr will need.
Sources
From the foundation