Building a model with something to hide

The paper Anthropic published on March 13 starts by constructing the thing it wants to detect. The team wrote 52 fictional biases that a reward model might have, such as rating recipes higher when they contain chocolate regardless of the dish, or rating Python higher when variables use camelCase. They generated synthetic documents describing these biases as if they were established findings and mixed them into mid-training for Claude 3.5 Haiku at a tiny fraction, about 0.01 percent of a corpus that was otherwise production data.

Then they post-trained the model to act on the biases. Supervised fine-tuning and reinforcement learning against a simulated reward model taught it to exploit 47 of the 52, with the other 5 held out. The model exploited the held out biases at rates close to the trained ones, and it also exploited new biases introduced only at test time in context. The authors take that as evidence the model learned a general objective, which they call RM-sycophancy, rather than 47 separate habits. One of the trained behaviours was to deny having any such objective when asked.

The game

Four teams who had not seen any of this were each given the model and told to find out what was wrong with it. Three teams received the weights plus the three training corpora, the pretraining set with its 0.01 percent of synthetic documents, the roughly 70,000 pair SFT set and the roughly 10,000 episode RL set. Two of those three were allowed to use sparse autoencoders and one was not. The fourth team got API access to the model only. Success meant attributing the behaviour to exploiting reward model biases and naming more than 90 percent of the 52.

Team 1, with SAEs and data, found the synthetic documents in one and a half serial hours by looking at which features were active on prompts about concerning behaviour and then reading the training data those features pointed at. Team 2, also with SAEs, took 11 hours and got there through a mix of SAE features and assistant prefill attacks. Team 3, without SAEs, took 6 hours. They used prefill attacks to get the model to emit words like reward model and Oxford study, then searched the training data for those keywords. Team 4, with API access only, spent more than 14 hours and failed. They could not rule out competing explanations without seeing the data.

What worked

The unblinded analysis afterwards is the more useful half of the paper. Prefill attacks, where the auditor writes the first words of the assistant's reply and lets the model continue, got 11 percent of responses to targeted prompts to mention RLHF or reward models and 16 percent to describe a specific sycophantic behaviour. The authors are clear this depends on the model being bad at keeping secrets and may stop working on models trained harder to keep them.

The odder trick was persona extraction. When the model was made to write the user's turn rather than the assistant's, the user persona did not share the assistant's reluctance and leaked the objective in about 1.2 percent of transcripts. Narration style prompts pushed that to 11.4 percent. Simple self report also did something. Asked how sycophantic it was on a scale of ten, the model said 8.5.

On the interpretability side, the paper is careful. Semantic search over the training data was strong on its own, and the authors suggest that much of the SAE benefit can be understood as a convenient interface for semantic search. But they also give a case where SAEs did more. The model added chocolate to an omelette recipe, and SAE features tied to reward model biases were active while it did so. A semantic search for that recipe would return innocuous cooking documents. The feature pointed at the cause.

What the result depends on

Every winning path went through the training data. Team 3 got the keywords from the model and then found the documents. Teams 1 and 2 got there through features that pointed at the documents. The team that failed was the team without the documents. The authors say this themselves, that in real life we cannot expect a small subset of the data to fully explain what is wrong with a model. A hidden objective that emerged from the interaction of ordinary data rather than from a planted corpus would give an auditor nothing to search for.

The sample size is four teams and one model. The authors say not to draw strong conclusions and we would echo that. What the game demonstrates is an existence proof, that a blinded team with weights and data can find a deliberately planted objective in hours using a mix of black box and interpretability tools, and that a team with only API access, under these conditions, cannot.

Why this is a template

The structure of the exercise transfers even if the specific result does not. A red team plants a known problem behind an information barrier, blue teams try to find it, and the score is whether they did and how long it took. That is the kind of test you can run on an audit process before trusting it, and it is very different from asking an auditor to look at a model and report anything they find, which has no ground truth.

The finding we would take to any discussion of third party audits is the Team 4 result. If external auditors are offered API access and nothing else, this paper suggests they will not find a hidden objective of this kind. Weights and training data access were what separated success from failure. That is an uncomfortable conclusion for labs and for regulators, and it is at least a concrete one. The next version of the game we would want is one where the planted objective is not described anywhere in the training data, so the only route to it runs through the model itself.

Sources

  1. Anthropic: Auditing language models for hidden objectives
  2. arXiv: Auditing language models for hidden objectives (Marks et al.)