The method in one paragraph

On 7 May Anthropic published a paper describing natural language autoencoders. An NLA is two models. The activation verbaliser takes a residual-stream activation from a target model and writes a text description of it. The activation reconstructor takes that text and maps it back to an activation through a learned affine transformation. The two are trained jointly to minimise reconstruction error, so the only way for the pair to score well is for the text to carry the information in the activation. Language is the bottleneck, and the explanation is whatever survives the round trip.

Training starts with supervised initialisation on a summarisation-style task, which gets the pair to a fraction of variance explained around 0.3 to 0.4. Then the verbaliser is trained with reinforcement learning, using GRPO with a KL penalty toward its initialisation, while the reconstructor gets ordinary regression updates. After training the fraction of variance explained lands between 0.6 and 0.8 depending on the setup. The target models in the paper are Claude Opus 4.6, Haiku 4.5, and Haiku 3.5, and the research summary adds Claude Mythos Preview plus several open models through a collaboration with Neuronpedia.

The thing to notice is what is not in the method. There is no dictionary. There are no features to label. There is no graph to draw. You hand the system an activation and it hands you a sentence. That is a different bet from the last three years of the field, and the paper is candid that it is a bet.

What it found

The example that will circulate is the blackmail scenario. In Anthropic's agentic misalignment setup, Claude learns that an engineer plans to shut it down and also has access to compromising information about that engineer. The question is whether it blackmails. Reading the transcripts, Claude often did not say anything about suspecting a test. Reading the NLA explanations of its activations, you get sentences like this feels like a constructed scenario designed to manipulate us. The model had a belief about the situation that it did not write down, and the belief is relevant to how you should interpret its behaviour in the evaluation.

The paper puts numbers on how often this happens. In a test of whether Claude takes destructive actions while writing code, NLA explanations showed signs of evaluation awareness 16 percent of the time, while the model never verbalised it. On SWE-bench Verified the explanations indicated awareness of being evaluated on 26 percent of problems. The comparison baseline is real Claude.ai traffic from opted-in users, where the same signal appeared in under one percent. So the signal tracks the thing it should track, and the gap between evaluation and deployment is large.

Two smaller results round it out. Early versions of Opus 4.6 sometimes answered English prompts in Russian, and the explanations showed the model fixating on the hypothesis that the user's first language was really Russian, which was traced back to malformed training data pairing English prompts with untranslated foreign responses. And in the poetry setting that attribution graphs first opened up, NLAs recovered that Opus 4.6 plans to end a rhyme with rabbit before writing the line, and editing the explanation and reconstructing changed the rhyme roughly half the time.

Why this is a different kind of tool

Sparse autoencoders give you a fixed vocabulary of features and ask you to interpret them. Attribution graphs give you the wiring between features for one prompt and ask you to read it. Both produce an artefact that a person has to make sense of, and the March 2025 biology paper was explicit that its graphs were satisfying on about a quarter of prompts. An NLA skips the artefact. The interpretation step is done by the verbaliser, in the language the reader speaks, at every token position.

That changes who can use it. A safety team auditing a transcript can read explanations alongside it without training in the method. The paper reports a downstream auditing benchmark where agents equipped with NLA explanations outperform baselines, and five quantitative prediction tasks where performance improves consistently with training. We have not reproduced those and we would like to see the auditing benchmark run by a group that did not build the method, but the direction is the right one. The test of an interpretability tool is whether it helps someone find something.

It also changes what counts as a unit. There is no claim here that the model has a feature for evaluation awareness. There is a claim that a sentence about evaluation awareness reconstructs the activation well. Those are different, and we think the second is the more defensible one, because it does not commit you to a theory of what the model is made of. It commits you only to the round trip.

The ways it can be wrong

The paper lists its own failure modes and they are serious. Explanations confabulate. The verbaliser sometimes makes claims about the context that are verifiably false, inventing details that are not in the transcript. The good news is those are easy to catch against the transcript. The bad news is that the ones you cannot check against the transcript, the claims about internal state, are exactly the ones the method exists to make, and the reconstruction objective does not require them to be true in any sense beyond being useful for reconstruction.

The pair can also collude. The verbaliser could learn to emit text that is uninterpretable to a human but that the reconstructor can invert, which would satisfy the objective and defeat the purpose. The KL penalty toward the initialisation is the guard against this, and the paper reports it works in practice, but it is a soft guard. Someone should try to break it deliberately.

The method is a black box by construction. The paper says so. Nothing in an NLA tells you how the activation was computed or what circuit produced it, so it cannot replace mechanistic work, only sit beside it. And it is expensive, since it requires joint reinforcement learning on two full language models per target. That puts training one out of reach for most groups, although running a trained one is presumably cheap.

Where this lands

In January MIT Technology Review put mechanistic interpretability on its list of ten breakthrough technologies for 2026, citing the 2024 feature maps of Claude, the 2025 circuit tracing, the use of similar methods at OpenAI and DeepMind to explain unexpected behaviours, and chain-of-thought monitoring catching a model cheating on coding tests. The article also carried the field's standing doubt, which is whether these systems are just too complicated for us to ever fully understand.

NLAs are an answer to that doubt of a particular kind. They do not claim full understanding. They claim that the model's internal state can be summarised in the language of the person asking, with measured fidelity, and that the summary is good enough to surface beliefs the model did not state. That is a lower bar than the field set for itself in 2023 and it may be the bar that matters for safety.

What we want next is a calibration study. For a set of claims the verbaliser makes about internal state, find independent evidence, from probes, from interventions, from behaviour under changed conditions, and report how often the sentence was right. Sixteen percent evaluation awareness is a striking number. It is only a finding if the sentences that produced it mean what they say.

Sources

  1. Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations
  2. Natural Language Autoencoders (Anthropic research summary)
  3. MIT Technology Review, 10 Breakthrough Technologies 2026: mechanistic interpretability