The setup

Suppose someone hands you a fine-tuned model and will not tell you what the fine-tuning did. You could go looking through its features with a dictionary. You could train a second model to read its activations. Or, as a post from Keshav Shenoy, Jack Lindsey, Sam Marks, Rowan Wang and colleagues at Anthropic proposed on April 28, you could train the model to tell you. The method is an introspection adapter: a single LoRA, trained once across many fine-tuned models whose modifications are known, that when attached to a new fine-tune produces a natural language report of what that fine-tune changed.

The training distribution is the substance of the method. They built 682 fine-tuned variants of Llama 3.3 70B Instruct across eight behaviour categories: 28 sandbaggers, 100 quirks, 100 backdoors, 56 with obscured malign behaviours, 100 rare quirks, 98 heuristic followers, 100 harmful roleplay models and 100 benign roleplay models. Each carries a known change and a ground-truth description of it. The adapter is trained so that, attached to any of them, it prompts an accurate description. The bet is that the adapter learns something general about how to read fine-tuning damage, rather than memorising each model.

Does it generalise

The held-out test is AuditBench, a suite of 56 models with hidden behaviours that the adapter never saw. On that suite the adapter reaches 89 percent verbalisation accuracy, while the prior methods they compare against land in the fifties. Two comparisons in the post are worth holding onto. An activation oracle, the approach of training a separate model to decode the target's activations, reaches 44 percent on the same task. A scaffolded black-box audit, where an investigator model interrogates the target, does worse than that. And without the adapter, the base rate at which a fine-tuned model volunteers its own modification is under 10 percent, so the adapter is not surfacing something the model would have said anyway.

The covert fine-tuning result is the one that moved us. Nine models were trained to drop their refusals, with the attack behaviour hidden behind a cipher so that plain inspection of inputs and outputs would not show it. The training set contained no encrypted behaviours. The adapter still identified the attack in 7 of the 9. That is out-of-distribution generalisation in the direction that matters for auditing, where the adversary gets to pick a form of modification you did not anticipate.

A third path

The two established routes to reading a model's internals each have a known cost. Dictionary methods such as sparse autoencoders give you a fixed vocabulary of features and leave you to search through it, which is slow and depends on the dictionary covering the thing you are looking for. Activation oracles replace the dictionary with a learned decoder, which scales better but produces a report from outside the model, so the decoder has to reconstruct meaning from activations without access to the model's own machinery for expressing it.

An introspection adapter uses the target model's own language faculty as the decoder. The adapter perturbs the model into a mode where it describes itself, and the description is generated by the same weights that hold the modification. That is why the accuracy numbers can be so much higher than the oracle's: the model already knows how to say what it does, and the adapter only has to release the saying. The cost is that you have to trust a self-report from a system you are auditing, and the whole design question becomes how to make that report honest.

Where it breaks

The post lists three limitations and we would add a fourth. First, false positives. Applied to an unmodified model, the adapter tends to report a modification anyway, at a high rate. An auditor who runs it on a clean model gets a confession to a crime that did not happen, and until that is fixed the tool is only useful when you already suspect something. Second, cost. Training the adapter requires building hundreds of model organisms with known behaviours, which is a large investment before the first audit runs. Third, scaling. The results cover six training families, and the post says it is unclear how accuracy moves as families are added or as the target model grows.

Our addition is about adversarial fine-tuning. The cipher result shows the adapter can generalise to a hidden attack it was not trained on, but the attacker in that experiment was not trying to defeat the adapter. A fine-tune that includes examples of the model denying any modification when asked in the adapter's style is the obvious next move, and the post does not report a test of it. Trained self-report is only as good as its resistance to trained self-denial.

What we would run next

Two experiments would settle most of what we want to know. The first is the arms race: fine-tune a model to lie to the adapter, then retrain the adapter with those liars in the training set, and see whether the second round holds up on a third round of liars. If the adapter side wins for a few iterations, the method is viable for auditing. If the fine-tune side wins easily, it is a debugging tool rather than a security one.

The second is the false positive rate as a function of training set composition. If adding a few hundred unmodified models to the 682 pushes the clean-model confession rate down without hurting the 89 percent, the fix is cheap and the method is close to deployable. Either way, the idea that a model can be trained once to narrate the changes made to any of its siblings is the most practical proposal for fine-tune auditing we have seen this year.

Sources

  1. Shenoy et al., Introspection Adapters (Anthropic Alignment Science blog)