Llama has a map and a calendar: linear probes for space and time
Gurnee and Tegmark fit ridge regression probes to Llama-2 activations and recovered latitude, longitude and dates for tens of thousands of real entities, with individual neurons that track the same directions. A reading note on what a probe result does and does not tell you about world models.
The setup
Wes Gurnee and Max Tegmark posted a paper on October 3 asking whether a language model's activations contain coordinates. They built six datasets of named entities with ground truth. Three are spatial, world places with 39,585 entries, US locations with 29,997 and New York City venues with 19,838, each labelled with latitude and longitude. Three are temporal, historical figures with 37,539 entries, works of art and entertainment with 31,321 and news headlines with 28,389, each labelled with a date. They ran the entity names through the Llama-2 family at 7B, 13B and 70B, took the residual stream at the last token of the name, and fit a linear ridge regression probe from that vector to the coordinate.
The probes work. On Llama-2-70B at the best layer the world dataset gives an R squared of 0.911, US locations 0.864, historical figures 0.835, entertainment 0.885 and headlines 0.746. New York City is the outlier at 0.359, which makes sense because the model has far less text about individual venues than about countries, and the coordinate range is tiny. Probe quality rises through the early layers and plateaus around the midpoint of the network.
Linear, and used
Two additional experiments give the result its weight. First, they swapped the linear probe for a one-layer MLP with 256 hidden units and got gains of only a point or two of R squared, so whatever structure exists is close to linear and a nonlinear decoder buys almost nothing. Second, probes trained on one entity type transfer to another, cities to landmarks for instance, which argues for a shared coordinate system rather than per-category lookup tables. The exception is entertainment, where the dates are bunched in recent decades and transfer suffers.
The neuron result is the part we find most convincing. They searched for individual neurons whose output weights align with the probe direction and found space neurons and time neurons whose activations track the true coordinate of the entity with no supervision at all. A probe can in principle read out something the model never uses. A neuron that fires in proportion to longitude is harder to explain away, because the model built that unit itself.
What prompting does
The robustness section is more nuanced than the abstract suggests. Asking the model explicitly for coordinates in the prompt made little to no difference to probe accuracy, which is what you would expect if the coordinates are stored on the entity rather than computed on demand. But random distracting tokens degraded the probes significantly, and so did changing capitalisation. The representation is there by default and it is fragile to context that pulls the last-token residual toward something else. We read that as the coordinate being one of several things superimposed in that vector, with the probe picking it out when the competition is quiet.
What a probe shows
The claim needs stating carefully. A linear probe with high R squared shows that the information is present and linearly accessible from the activations at that layer. It does not by itself show that the model reads it in that form when answering a question, and it does not show that the model does anything like reasoning over a map. The paper says as much, describing what it found as basic ingredients of a world model and stating that the true extent and structure remain unclear.
The gap between present and used is where we would want the next experiment to go. The obvious test is causal. Take the probe direction, shift an activation along it so that Paris moves to the coordinates of Madrid, and see whether the model's downstream answers about neighbours, time zones or travel change accordingly. If they do, the map is doing work. If they do not, the coordinate is a correlate the model stores and mostly ignores. Marks and Tegmark's companion paper on true and false statements, posted a week later, runs exactly that kind of intervention for a truth direction and finds that simple difference-in-means directions have stronger causal effects than fancier probes, which is encouraging for the same approach here.
Why it matters for the world model argument
The word world model gets used loosely, and this paper is a good chance to tighten it. There is a weak claim, that the model's internal state encodes facts about the world in a structured and reusable way, and a strong claim, that the model simulates the world to make predictions. The evidence here supports the weak claim clearly for two specific quantities. It says nothing directly about the strong one.
That is still more than a model that had memorised strings would give you. A lookup table of city names does not need to place them on a shared plane with a linear geometry, and it certainly does not need a neuron for longitude. Whatever Llama is doing with these coordinates, it built a coordinate system to do it. We would like to see the same method applied to quantities that are harder to read off training text, like population or elevation, to find where the map stops.
Sources
From the foundation