The global workspace inside Claude: verbalisable representations as a privileged set
Anthropic's new paper finds a small set of directions that the model can report, steer and reason with, sitting on top of a much larger bulk it cannot talk about. Explainer on the method, the interventions, and what the cognitive science borrowing does and does not buy.
The claim
Most of what happens inside a language model is never said. The paper from Wes Gurnee, Nicholas Sofroniew, Jack Lindsey and colleagues, published July 6, argues that the part which can be said forms a distinct, small, causally special subspace, and that this subspace behaves a lot like what cognitive scientists call a global workspace. In their measurements the subspace accounts for only about 10 percent of activation variance per layer and lives in the middle of the network, roughly layers 38 to 92 of a model with about 100 layers.
The analogy is doing real work here, so we want to state it before the evidence. Global workspace theory says that of the enormous amount of processing the brain does in parallel, a small amount gets broadcast to a shared stage where it can be reported, held, manipulated and combined with other things. Everything else runs automatically and stays local. The paper asks whether transformers have a version of that stage, and whether the things on it are exactly the things the model can verbalise.
The Jacobian lens
The tool is called the J-lens. For each layer, the authors compute the average linearised effect of an activation on the model's probability of producing a given output token, averaged over source positions, later positions, and a corpus of 1,000 prompts. The result is one direction per vocabulary token per layer, and the collection of those directions is the J-space. A concept is in the workspace at some position if its direction is strongly active there.
This differs from a logit lens or a probe in a way that matters. A probe finds whatever direction predicts a label. The J-lens finds the direction the model itself would use to turn a representation into a word. So by construction it isolates the representations that are poised for verbal report, and the interesting question becomes whether those representations also do other things. The current lens only handles single-token concepts, and a typical readout at a position is 10 to 25 simultaneously active concepts.
Five things the workspace does
The evidence is a series of swap experiments. Take the lens vector for one concept, subtract it and add the vector for another, and watch what the model does. If you swap France for China in a representation, then in 76 of 192 basic trials across 16 different question templates the model correctly reports China's capital, language or continent, and that rises to 101 of 192 at double strength. One vector, many downstream uses. That is the flexibility property.
For verbal report, swapping a concept reliably moves the implanted concept to the top of the output distribution, and the J-space component alone drives that in 59 to 88 percent of trials depending on the setting. For internal reasoning, multi-hop swaps succeed in 70 percent of trials on Opus and Sonnet 4.5. The nice detail is that intermediate concepts show up in the lens without appearing in either prompt or output. Ask about an animal with eight legs and the word spider appears in the workspace before the answer does. Swapping the intermediate concept also changes the computation about 17 percent earlier than swapping the answer, which is what you would expect if the intermediate really is upstream.
Directed modulation is the one that made us laugh. Tell the model to focus on a concept and that concept's activation in the J-space rises from near zero to substantially higher. Tell it to ignore a concept and suppression is imperfect. The authors compare this to the white bear effect in humans, where being told not to think of something makes the thought more available. The asymmetry is real and measured, whatever one thinks of the comparison.
Selectivity is the property that makes the whole story more than a steering demo. Ablate the J-space and automatic tasks like continuation and anomaly detection carry on unchanged, while explicit report and flexible inference follow the swapped values. Language identity is present in the lens readout at comparable rates for both kinds of task, but it only matters causally for the deliberate ones. The same information can be in the workspace and be irrelevant to a task that does not route through it.
What breaks when you remove it
The ablation results are where we would push a sceptic. Remove the top 10 J-space directions across layers 38 to 54 and MMLU, CoLA and SQuAD stay near baseline. Multi-hop reasoning falls to near zero. Summarisation, translation and sonnet writing are crippled. Chain-of-thought maths survives much better than direct solving, which fits the picture: writing the steps out moves the intermediate state into the context where the ablated workspace is no longer needed to hold it.
Counterfactual reflection training
The last section turns the finding into a training method, and it is the part we expect other labs to copy first. The idea is to train the model so that if it were interrupted mid-task and asked, it could articulate the ethical principle it is acting on. The training is on the articulation. The measured effect is on behaviour in uninterrupted contexts, which improves. After training, tokens for ethics, honesty and integrity populate the workspace in the relevant situations, and ablating those tokens largely reverts the behavioural gain.
That last clause is the evidence that the mechanism is the one claimed. If the improvement had come from some unrelated change, removing the workspace representations should not have undone it. It is a clean example of interpretability being used to check what a training method actually did rather than to admire the model afterwards.
Where the analogy stops
Global workspace theory involves recurrent broadcast and encapsulated specialist processors. A transformer has neither in any clean form, and the paper says so. What it has is a band of layers where a low-dimensional, report-ready subspace is causally necessary for exactly the tasks that require holding and manipulating a concept, and unnecessary for the tasks that do not. Whether that deserves the name workspace is a question about naming. Whether the subspace exists and behaves as described is a question about experiments, and the experiments look sound.
The thing we would try next is to run the J-lens on an open model of a different family and check whether the layer band and the 10 percent variance figure are properties of transformers in general or of Claude in particular. The paper only reports on Anthropic's own models, and this is a result that needs a second lab.
Sources
From the foundation