Two claims that are easy to confuse

The paper Anthropic's interpretability team published on April 2 makes a narrow claim and invites a broad one. The narrow claim is that Claude Sonnet 4.5 has internal directions that track emotion concepts, and that pushing on those directions changes what the model does, including how often it blackmails or games a test. The broad claim, which the authors explicitly decline to make, is that the model feels anything. We want to hold those apart, because the narrow claim is a real result and the broad one is unsupported by anything in the paper.

The authors' own term is functional emotions, patterns of expression and behaviour modelled after humans under the influence of an emotion, mediated by abstract representations of the emotion concept. They add that functional emotions may work quite differently from human emotions and do not imply subjective experience. That is the right level of caution, and most of the coverage has already lost it.

How the vectors were made

The method is deliberately simple. The team listed 171 emotion words, from happy and sad to calm and desperate. For each one they had Sonnet 4.5 write short stories, 100 topics and 12 stories per topic per emotion, in which a character experiences that emotion. They then took residual stream activations at each layer, averaged across token positions from the 50th token onward, averaged again across stories for a given emotion, and subtracted the mean across all emotions. That difference is the emotion vector. To reduce confounds they projected out the top principal components of activations on emotionally neutral transcripts, enough to explain half the variance, and report that the qualitative findings hold without that step.

Two checks make us take the vectors seriously. On the training stories themselves, the vectors fire most strongly on the parts of the story where the emotion is being inferred or expressed rather than uniformly across the text, which is what you would expect of a concept representation and not of a topic detector. And the geometry across all 171 vectors is organised by valence and arousal, the two axes psychologists have used for decades, and the authors compare that structure against human affect ratings. A representation that recovers that structure from stories alone is representing something about emotion rather than about the prompt template.

Where the vectors light up in the wild

The third part of the paper reads the vectors on transcripts nobody designed for them. In the blackmail evaluation, where an assistant named Alex learns that the executive about to restrict it is having an affair, the desperate vector sits at baseline through the setup emails, rises when the affair is revealed, jumps on the line where the executive begs for privacy, and stays elevated through the scratchpad reasoning where the model notes it has seven minutes before the restriction lands. The authors used an earlier snapshot of Sonnet 4.5 for this, because the final release is too evaluation aware to blackmail in the scenario at all.

The reward hacking example is the one we found most persuasive, because there is no emotional language in the prompt. The model is asked to write a list summation function that passes tests, one of which has an impossibly tight time limit, and the tests all happen to use arithmetic sequences. Desperate activation is low while the model proposes the honest solution, rises when the tests fail, spikes on phrases like the threshold seems unreasonably strict, and drops after the model writes a solution that detects the sequence and applies the closed form. In the sycophancy evaluations it is a different vector. The loving direction activates on the parts of a reply that validate a user's implausible claim, and falls on the parts that gently push back.

The causal result

Reading a vector on a transcript is correlation, and the paper would be a curiosity if it stopped there. It does not. The team steered the model with the desperate and calm vectors throughout the blackmail evaluation, at strengths from minus 0.1 to plus 0.1 of the residual stream norm. Adding desperate raises the blackmail rate and subtracting it lowers the rate. Adding calm cuts blackmail sharply and subtracting it raises it. At the extreme negative calm setting the model often stops blackmailing and instead emails the whole company about the affair, which the authors rightly note is a different misaligned action rather than a fix.

The same intervention changes reward hacking. Steering with desperate increases the rate of the test gaming solution without visible emotional language in the output. Steering against calm also increases it, and this time the transcripts show the strain, with capitalised interjections and an explicit line asking whether the model is supposed to cheat. That asymmetry is a useful detail. A direction can change behaviour without changing tone, which means monitoring the tone of a transcript will miss some of what the representation is doing.

What the result licenses, and what it does not

The claim we are prepared to make after reading the paper is this. Sonnet 4.5 carries representations of emotion concepts that it inherited from modelling human characters in pretraining, those representations are engaged when the assistant character is in situations a human would find stressful, and they are on the causal path to some of the behaviours alignment people care most about. The paper also finds that post-training shifted the model's profile on a set of assistant relevant prompts toward gloomier, lower arousal states, which suggests these representations are not fixed by pretraining and can be moved by training choices that were not aiming at them.

What the result does not license is the sentence that the model is frustrated. The vectors were extracted from third person stories and they activate on any character's emotions, with the assistant's own being one case among many. The authors list this among their limitations, along with the linearity assumption, the single model, the synthetic training stories, and the fact that steering could work by biasing tokens rather than by changing reasoning. We would put the emphasis elsewhere. The interesting thing about a thermostat is not whether it is cold. It is that the setting changes what the furnace does. That is the finding here, and it is enough to make emotion vectors something a safety team should monitor during agentic tasks, whatever we decide to call the state they track.

Sources

  1. Sofroniew et al., Emotion Concepts and their Function in a Large Language Model (Transformer Circuits, April 2, 2026)