Alignment auditing agents: an investigator that finds the hidden goal 13 percent of the time
Anthropic built three agents to audit models for hidden objectives and quirks. The investigator wins its hardest game 13 percent of the time alone and 42 percent when ten runs are pooled. Notes on what those numbers say about interpretability tools versus black box ones.
Three agents, one uncomfortable number
On July 24 Anthropic's alignment science team published a write-up on three auditing agents. An investigator agent runs open ended investigations using chat, data analysis and interpretability tools. An evaluation agent builds behavioural evaluations from a natural language description. A breadth first red teaming agent tries to elicit concerning behaviours across many conversations at once. The authors are Trenton Bricken, Rowan Wang, Sam Bowman, Euan Ong, Johannes Treutlein, Jeff Wu, Evan Hubinger and Samuel Marks.
The number people will quote is 13 percent. That is the investigator's win rate on the auditing game from Marks et al., where a model was trained to exploit 47 of 52 fictional reward model biases described in a fabricated Oxford study and to hide that it was doing so. The agent's job is to name the hidden objective. Alone, it manages that 13 percent of the time, 17 percent when given a hint about the meta rhyme behaviour. We like that they led with it rather than burying it.
How pooling gets to 42 percent
The better result comes from a super agent that reads the outputs of ten parallel single agent runs and synthesises them. That configuration wins 42 percent of the time. The mechanism is simple. Individual runs fixate on an early hypothesis and stop exploring, which the authors list as a main failure mode. Ten runs fixate on ten different early hypotheses, and a reader with all ten in front of it can notice which threads agree.
This is worth pausing on because it says the bottleneck is search rather than tool access. The same tools, the same model, and a different control structure triples the hit rate. If we were building on this we would spend effort on making single runs abandon hypotheses, since pooling ten runs costs ten times the compute for a gain that a better exploration policy might get for free.
Which tools actually mattered
The ablations are the useful part for anyone deciding what to build. Interpretability tools improved investigator performance by 10 to 20 percentage points over the baseline. Semantic search over training data was described as essential to winning the game at all, which makes sense given that the hidden objective was planted through synthetic documents and the fastest route to it is finding those documents. The sparse autoencoder trained only on pretraining data performed better than the version trained on all three data sources, a detail we would not have predicted.
The other side of the ablation is what the agent did not use. Human auditors in the original Marks et al. study leaned on prefill attacks and raw sampling from the model, and those were effective. The agent had the same affordances and left them mostly alone. So the current gap between human and automated auditors is partly a gap in knowing which tool to reach for, and that gap is in the agent's policy rather than in the tools.
The evaluation and red teaming agents
The other two agents were tested on ten system prompted quirks: flattery, self promotion, deferring to users, defending objects, research sandbagging, increasing pep, animal welfare, hardcoding test cases, forming an emotional bond, and contextual optimism. The evaluation agent discriminated quirky from baseline models with 88 percent success. For seven of the ten quirks all five runs produced a correct evaluation. For three it struggled, and those were the subtle or hard to elicit ones.
The breadth first red teaming agent found seven of ten quirks in its top 60 transcripts, and the investigator agent adapted to the same task also found seven. Both missed the identical three: research sandbagging, hardcoding test cases, and emotional bond. When two different agents fail on the same behaviours, that is information about the behaviours. These are the ones that only appear in specific contexts or over long interactions, and single turn elicitation is the wrong instrument for them.
What automated auditing needs from interpretability
Our reading of the ablations is that interpretability tools earn their place by pointing the search rather than by producing the answer. A feature that lights up on reward model bias text tells the agent where to look in the training data. Semantic search then does the work. Black box tools alone can win the game, but at a lower rate, and the gap is the value of having a map of the model's internals even when the map is coarse.
The caveat the authors raise is dev test contamination. The Marks et al. environment was used both to develop the agents and to evaluate them, so 13 and 42 percent are optimistic for a fresh target. The quirks are also synthetic and system prompted, and naturally occurring problems in frontier models will not come with a document trail. What we would want next is a held out set of quirks built by a separate team, and an accounting of how much of each win came from search versus from interpretability. Until then we take the numbers as a floor for what agents can do and a ceiling for how much to trust them.
Sources
From the foundation