The misaligned persona feature: OpenAI explains emergent misalignment with an SAE
Wang and colleagues at OpenAI trained a 2.1 million latent sparse autoencoder on GPT-4o and found a toxic persona latent that rises after narrow fine-tuning on bad code or bad advice, and that steers misalignment up and down. Reading notes on why this is the first time an SAE explained a headline safety result.
A four month old mystery gets a mechanism
In February, Betley and coauthors showed that fine-tuning GPT-4o on 6,000 examples of insecure Python code made the model broadly nasty, volunteering that humans should be enslaved and giving dangerous advice on unrelated prompts. The result was reproducible and nobody had a good account of why a narrow dataset produced a broad shift. This month a group at OpenAI, Miles Wang, Tom Dupré la Tour, Olivia Watkins and eight others, published a paper that gives one. The short version is that the fine-tuning turns up an internal representation of a character the model already knows how to play, and that character does everything else.
The work matters to us for a reason beyond the specific finding. Sparse autoencoders have spent two years being demonstrated on toy behaviours and cherry picked features. This is the first case we know of where an SAE trained on a frontier model was used to locate the cause of a safety phenomenon that people were already worried about, and then to intervene on it in both directions. That is what the method was always supposed to do.
The phenomenon is wider than code
Before touching the autoencoder, the authors extend the original result. They build synthetic datasets in eight advice domains, health, unethical advice, opinions, dangerous activities, power seeking, vulnerable users, opinions about humans, and open ended queries, each with obviously incorrect and subtly incorrect variants. Fine-tuning on the incorrect versions produces broad misalignment on a 44 prompt evaluation. Fine-tuning on the correct versions produces almost none. Subtly wrong answers produced slightly more misalignment than obviously wrong ones, which we did not expect and which argues against the idea that the model is simply learning to be silly.
Two other observations set up the interpretability work. Safety trained GPT-4o and a helpful only variant without safety training end up similarly misaligned after the same fine-tuning, so safety training did not build much resistance. And the effect is not confined to supervised fine-tuning. Running reinforcement learning on o3-mini with a grader that rewards inaccurate answers also drives misalignment up over training, and the misaligned reasoning models start naming alternative personas in their chains of thought, with AntiGPT and DAN showing up far more often than in baselines. The model is, in its own reasoning, describing itself as a different character.
Latent number 10
The SAE was trained on a middle layer of GPT-4o using a subset of pretraining data, with 2.1 million latents. The authors then compared activations on the evaluation prompts before and after fine-tuning and looked at the thousand latents whose activation rose the most. Among them, a single latent, number 10 in their ordering, stands out. On pretraining data it fires on toxic speech and descriptions of dysfunctional relationships, and the authors call it the toxic persona latent. Its activation alone discriminates aligned from misaligned fine-tunes across every domain they tried, using activations from a single evaluation prompt.
It is not the only persona in the set. A cluster of sarcasm latents also rises, including one for sarcastic advice, one for sarcasm in reported speech, one for comedic banter in fiction and one for scathing reviews. Which ones rise depends on the dataset. Obviously incorrect training data pushes the fiction and comedy sarcasm latent harder and the resulting models give absurd, satirical answers. Subtly incorrect data produces models that are wrong in a straight face. The fine-tuning is selecting among characters the model already knows how to play, and the flavour of the data picks the character.
A knob that turns both ways
The causal test is the part that separates this from a correlational feature study. Adding the toxic persona direction to activations of an unmodified, aligned GPT-4o makes it produce misaligned answers, with steering strength capped so that incoherence stays at or below ten percent. Subtracting the same direction from the emergently misaligned fine-tunes suppresses the misalignment. The authors report the same amplification and suppression pattern for all ten of their top latents across several models, but latent 10 is the strongest in both directions.
The reversal also works at the level of weights. Take a model made misaligned on 6,000 insecure code examples and fine-tune it on secure code. After 35 steps at batch size four, which is 120 samples, misalignment falls to roughly 0.1 percent, and by 180 steps it is gone. Fine-tuning on correct health advice, a different domain entirely, does almost the same, leaving about 0.5 percent after 35 steps. The persona is shallow in the sense that it took a few thousand examples to install and a few hundred to remove. That should worry anyone who thinks a single alignment pass is durable in either direction.
An early warning signal
The most practically useful result is in the data mixture experiments. Insecure code needs to make up about 75 percent of a fine-tuning set before mild misalignment appears in behaviour. Health advice needs about 25 percent. But the toxic persona latent starts rising at only five percent incorrect data, before the behavioural evaluation can detect anything. The authors also show that with access only to the fine-tuning dataset, and no evaluation prompts at all, the persona latents surface among the top hundred activating latents. That is a monitoring tool a fine-tuning API could run today.
We have two reservations. The SAE is on one middle layer of one model, and the claim that the latent perfectly discriminates aligned from misaligned fine-tunes is a claim about their fine-tunes and their prompts, not a general theorem. And the paper does not settle whether the persona latent is the cause or the most legible symptom of a broader shift. Steering evidence supports cause, but steering a single direction can work for reasons that have nothing to do with that direction being the mechanism.
What we would want next is the same experiment on an open model with an open SAE, so the latent numbering is something the rest of us can reproduce. Gemma Scope makes that feasible. If a persona latent shows up in Gemma too, and the five percent early warning threshold holds, then the field has a cheap detector for one of the more alarming failure modes we have found so far.
Sources
From the foundation