Othello-GPT's board is linear after all
Li et al. found a nonlinear world model in a transformer trained on Othello moves. Nanda, Lee and Wattenberg show that if you ask the probe about mine versus theirs instead of black versus white, the same board is linear and steerable with vector arithmetic.
The original result
Last year Kenneth Li, Martin Wattenberg and colleagues trained an 8-layer, 512-dimensional GPT on synthetic Othello games, feeding it nothing but move sequences and asking it to predict legal next moves. Then they asked whether the model had built an internal picture of the board. Their answer was yes, with a caveat that got a lot of attention. Two-layer MLP probes could read the board state out of the residual stream with 1.7 percent error. Linear probes could not, landing around 20.4 percent error. They concluded the world model was nonlinear.
That conclusion mattered beyond Othello. A lot of interpretability work rests on the hope that features are linear directions in activation space, because linear things are the ones we know how to find, ablate and steer. A clean synthetic task where the representation turned out to need an MLP was evidence against that hope. We cited it that way ourselves.
Ask a different question
Neel Nanda's replication, now written up as a paper with Andrew Lee and Wattenberg, changes one thing about the probe. The Li et al. probes asked whether a square held a black piece, a white piece or nothing. Nanda's probes ask whether a square holds a piece belonging to the player about to move, a piece belonging to the opponent, or nothing. Mine, theirs, empty.
The reason that works is that the model has no idea which colour it is. Othello alternates turns, and the network sees the same kind of token on every move. From the model's point of view the natural variable is whose turn it is, not what colour the pieces are. A black piece on move seven is mine, and the same black piece on move eight is theirs. A linear probe for black must therefore flip sign on every move, and no fixed direction can do that. A linear probe for mine does not need to flip at all.
The way Nanda found it was almost mechanical. He trained separate linear probes on odd moves, where black is to play, and on even moves, where white is to play. Each was nearly perfect. The two probe directions were close to negations of each other, and one could be transferred to the other by flipping the sign. Once you see that, the mine and theirs framing writes itself.
Steering the model with arithmetic
With a linear representation in hand, intervention becomes simple. The paper takes the probe direction for a square, negates the model's activation along that direction, and checks whether the model now predicts moves that are legal on the edited board. It does. The authors are candid that this part is less rigorous than the probing, and it rests on case studies across a few hyperparameter settings rather than a full sweep. But the Li et al. interventions needed gradient descent to edit the board. This one is one subtraction.
Probe accuracy is highest around layers four and six, and probes trained on one layer transfer zero-shot to nearby layers, which suggests the direction is stable through the middle of the network rather than being re-encoded each block. Corners are the weak spot. Accuracy drops there, and the paper does not fully explain why.
What the replication actually says
We want to be careful about the lesson. The Li et al. result was not wrong. Their probes measured what they measured, and a two-layer MLP does recover black and white with low error while a linear probe does not. What changed is the target variable. The model's representation was linear in one basis and nonlinear in a closely related basis, and the first team happened to pick the second.
That is uncomfortable for the field, because it means a negative result about linearity is only as strong as the set of probe targets someone thought to try. The board looked nonlinear because black versus white is the way humans describe Othello. The model had no reason to share that description. Any time a probe fails, the honest next step is to ask whether the concept was specified in the model's terms or in ours, and there is no general procedure for doing that.
Nanda's own summary is that this is moderate evidence for the linear representation hypothesis, and we would put it the same way. One synthetic game, one architecture, one clever reparametrisation. It does move us, because the failure mode it exposes, a sign flip driven by turn order, is exactly the kind of thing that could be hiding behind nonlinear results elsewhere.
A check worth running elsewhere
The cheap experiment this suggests is to revisit published nonlinear probe results and look for a hidden alternation. Anywhere a model processes a sequence with a role that swaps, such as speaker turns in dialogue or the two sides of a negotiation, a probe trained on absolute labels will struggle in the same way. Retraining with relative labels takes an afternoon.
The harder question is what to do when there is no obvious relative frame to try. We do not have an answer. What we take from this paper is that the probe target is a hypothesis about the model, and a failed linear probe rejects the hypothesis rather than the model.
Sources
From the foundation