Concept injection: Claude notices its own thoughts about 20 percent of the time
Jack Lindsey injected concept vectors into Claude's activations and asked whether the model noticed. Opus 4.1 did on about a fifth of trials at the best layer and strength. Why this is an interpretability result about a causal intervention, and what the failures say.
The intervention
Jack Lindsey at Anthropic posted a paper yesterday asking whether a language model can report on its own internal state, and the method is what makes the question answerable. The technique is concept injection. You take an activation pattern that represents a concept, add it into the model's residual stream at some layer, and then ask the model whether it notices anything unusual. Because you put the concept there, you know the ground truth. A self-report can be scored as right or wrong.
That is the difference between this and asking a model to describe its reasoning. When a model explains a chain of thought, there is no ground truth to check against, so a plausible explanation and a true one look the same. Here the experimenter controls the internal state, and the report either matches it or does not.
What the models did
At the best injection strength and layer, Claude Opus 4.1 correctly noticed and named the injected concept on about 20 percent of trials. Opus 4 and 4.1 were the strongest, older production models were weaker, and Sonnet and Haiku variants were also tested. Lindsey's example in the VentureBeat interview is that the model does not simply repeat the injected word. It reports that it is thinking about betrayal, which means it is describing the state rather than being pushed to emit the token.
There are four experiments. The first is the injected thought detection above. The second checks that a model can transcribe real input text while separately reporting an injected thought, so the two channels are distinguishable. The third is the prefill experiment. If you force the model to emit an unnatural word, bread, and then ask whether it meant to, the model disavows the output. If you retroactively inject the activation for bread before the output, the model claims it. The fourth asks the model to think about a concept, or to avoid thinking about it, and measures whether the internal representation moves.
Why this is an interpretability paper
The paper states that it does not seek to address whether AI systems have human-like self-awareness or subjective experience, and we think that framing is right on the merits as well as on caution. The result is a fact about a causal relationship. An intervention on activations produces a change in a verbal report that tracks the intervention, some of the time. That is a claim about mechanism. It says something inside the model reads other things inside the model and can route the reading to the output.
Read that way, the finding belongs next to activation steering and probing rather than next to philosophy. Steering showed that adding a vector changes behaviour. Probing showed that internal states are linearly readable by an external classifier. This paper shows that the model itself can sometimes act as the classifier. The 20 percent is the accuracy of that internal probe under the best conditions found.
What the failures teach
The failure modes are listed and they are instructive. Weak injections are missed. The model can be influenced by an injection while denying that it detected anything, which is the worst case for anyone hoping to use self-report as a monitor. High steering strengths produce what the paper calls brain damage, incoherent output where any report is noise. Recognition sometimes comes late. And some model variants produce false positives, reporting an injected thought when nothing was injected.
Lindsey's own summary in the interview is that right now you should not trust models when they tell you about their reasoning, and that the models are getting smarter faster than we are getting better at understanding them. Those two statements, from the author of the most positive introspection result to date, set the ceiling on how far anyone should push the 20 percent.
What we would try
The obvious next experiment is to train for it. If introspective accuracy is 20 percent emergent, it may be far higher with a small amount of supervised practice on injected concepts, and then the question becomes whether the trained ability generalises to concepts and layers it was not trained on. The second is to test whether a model can detect an injection that changes its behaviour in a way it would otherwise hide, since that is the monitoring use case people actually want.
And the false positive rate needs its own study. A model that reports thoughts it does not have is confabulating, and confabulation in self-report is exactly what this method exists to catch. If the false positive rate rises with model capability, the 20 percent number gets harder to interpret.
Sources
From the foundation