Grokking, reverse engineered: the Fourier circuit inside modular addition
Nanda and colleagues opened up a one-layer transformer that groks modular addition and found trig identities. The progress measures they built from that circuit move long before test loss does.
The setup
Grokking is the name for a training curve that looks broken. A small network memorises its training set, sits at chance on held-out data for thousands of steps, and then, with nothing in the training loss to warn you, generalises. The standard demonstration is modular arithmetic. A paper out this week from Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith and Jacob Steinhardt takes that demonstration apart and explains, weight by weight, what the network is doing on either side of the jump.
The model is a one-layer ReLU transformer with embedding dimension 128, four attention heads of width 32, and 512 MLP neurons, trained to output a plus b mod 113. No layer norm, untied embedding and unembedding. Training uses 30 percent of the 113 by 113 possible input pairs, full batch AdamW at learning rate 0.001, weight decay 1, for 40,000 epochs. Train loss drops early. Test accuracy jumps at around epoch 10,000.
The algorithm the network learned
The trick is to look at the embeddings in the Fourier basis over the integers mod 113. Almost all the mass sits on five frequencies, k in the set 14, 35, 41, 42 and 52. The embedding layer maps each input to sines and cosines at those frequencies. The attention and MLP layers then multiply them together, and by the product-to-sum identities that turns sin and cos of wa and wb into sin and cos of w times a plus b. The network has converted addition into rotation around a circle.
The output side uses one more identity. The logit for a candidate answer c is a weighted sum of cos of w times a plus b minus c, computed as cos of w times a plus b times cos of wc plus sin of w times a plus b times sin of wc. For each key frequency that term peaks when c equals a plus b mod 113. With five frequencies summed, the peaks line up at the correct answer and cancel elsewhere. That is the whole algorithm, and the paper calls it Fourier multiplication.
The evidence is more than a pretty plot. The neuron-to-logit map is rank 10 and is approximated by the sine and cosine directions at the five frequencies with 0.55 percent residual error. 84.6 percent of the MLP neurons are well approximated by degree-two polynomials of a single frequency's sinusoids, with more than 85 percent of variance explained. Ablate the five key frequencies and the network drops to chance. Ablate the other 95 percent of Fourier components and loss goes down by about 70 percent. The rest of the network is noise the algorithm has to fight through.
Progress measures, and what happens before the jump
Once you know the circuit you can measure how much of it exists at every checkpoint, and that is the paper's real contribution. Restricted loss keeps only the 20 Fourier terms belonging to the key frequencies and throws away everything else. If the generalising circuit is present, restricted loss is low. Excluded loss does the opposite, removing only the key frequencies from the logits. If the network is still leaning on memorisation, excluded loss stays low, because memorisation lives in the other components.
Plotted together these two measures split training into three phases. From epoch 0 to about 1,400, train loss falls while test and restricted loss stay high. That is memorisation. From about 1,400 to 9,400, excluded loss rises and restricted loss falls while train and test loss barely move. The circuit is being built underneath a flat curve. From 9,400 to about 14,000, test accuracy jumps, weight norm drops sharply, and the Gini coefficient of the Fourier component norms rises, meaning the spectrum gets sparse. The memorising components are being cleaned out.
So the phase transition in test loss comes at the end of a long process rather than at its start. Circuit formation took roughly 8,000 epochs and was fully visible in restricted loss the whole time. The sudden part is only the cleanup, and cleanup is what weight decay does once the circuit is good enough that the memorising weights are pure cost. The authors frame grokking as gradual amplification of a structured mechanism followed by removal of the memorising components, and the measures back that up.
Why this sets a bar
Mechanistic interpretability has produced a lot of partial stories. A head that seems to do induction. A neuron that fires on French. This paper is different in kind because the explanation is complete enough to be tested by ablation in the basis it predicts, and because it makes a quantitative prediction about training dynamics that turned out true. If your reverse engineering is right, you should be able to build a progress measure from it and watch it move before the metric everyone else is looking at. They did, and it did.
The honest caveat is scale. One layer, 128 dimensions, a task with a known closed-form solution. The paper checks that the story holds across seeds and a few sizes, but the entire circuit fits in a handful of frequencies because the task is periodic. There is no reason to expect real language models to be this clean, and the authors do not claim otherwise. What transfers is the method. Guess the mechanism, express it in the right basis, ablate everything else, and turn the mechanism into a measurement.
What we would try next
The obvious experiment is to intervene on the phases. If restricted loss really tracks circuit formation, then adding a term to the training loss that rewards low restricted loss should shorten the flat region, and removing weight decay should stall cleanup indefinitely while leaving the circuit half-built. Either result would tighten the causal story from correlation to control.
The more ambitious one is to find any task in a real model where a progress measure of this kind can be built. Even a crude one, a few directions in the residual stream that a circuit is known to use, would let us ask whether capabilities in large models also form gradually and appear suddenly. If they do, a lot of what looks like emergence is a measurement artefact, and that is a claim worth the effort of checking.
Sources
From the foundation