The model organism lottery: our test subjects may be too easy
Fifty-four small models with the same implanted quirks, built seven different ways, gave interpretability tools scores that varied by up to twenty times. Models trained the realistic way were usually the hardest to read. On what that does to the auditing results we have been quoting.
How we test interpretability tools
When someone wants to know whether an auditing technique can find a hidden behaviour in a language model, they build a model organism. Take a small open model, fine-tune it to hold a false belief or a strange preference, then check whether the technique recovers the quirk without being told what it is. The field's recent optimism, including the blind auditing games where teams found a hidden objective and the activation oracle results that pulled secret words out of Taboo models, rests on organisms built this way.
A paper from LASR Labs and Cambridge, posted July 1 and presented at the ICML mechanistic interpretability workshop, asks a question that in hindsight should have been asked earlier. Does the result depend on how the organism was made? Andrzej Szablewski, Gabriel Konar-Steenberg, Raffaello Fornasiere, Nikita Menon, and Stefan Heimersheim built a suite of 54 organisms to find out, and the answer is yes, by a lot.
The suite
Three quirk families, all benign. False facts about baking cakes, borrowed from earlier work. A preference for Italian food. A fixation on submarines in military contexts. Each family was instilled into OLMo2-1B using seven training methods. Standard transcript distillation and synthetic document fine-tuning, each with and without a one-to-one mix of unrelated C4 data. Post-hoc DPO, with and without mixing. And a new integrated method that folds the quirk data into OLMo's original open post-training DPO phase, which is meant to look like how a behaviour would actually arrive during a real training run.
Two of the families were also built on gemma-3-1b-it to test architecture, and the cake family was replicated across three data-ordering seeds to test noise. The detail that makes the study work is expression matching. Within each family every variant was tuned to express the quirk at the same rate and to leak it into unrelated contexts as little as possible, so differences in interpretability cannot be explained by one model simply having the quirk more strongly.
Four techniques were then run across the suite. Activation oracles, which are models trained to describe another model's activations in natural language. Activation steering along the average difference between the organism and the original model. A logit lens method that projects activation differences to the vocabulary and scores how many top tokens relate to the quirk. And sparse autoencoders.
What varied
After dropping variants that fell below a per-family noise floor, the ratio between the most and least interpretable variant in a family ranged from 1.2 times to 20.4 times, depending on family and technique. The ranking of training methods did not generalise across families, techniques, or architectures. It was, however, stable across training seeds, so the differences are properties of the recipe and not of run-to-run noise.
The integrated DPO variant, the one built to resemble real training, scored lowest or second lowest on activation oracles and steering in every family and on both architectures. The logit lens gave one counterexample on the Italian food family. So the technique that produces the most realistic organism also produces the hardest one to read, which is the outcome you would least want if you were hoping post-hoc organisms were representative.
Data mixing did not behave as expected either. Earlier work had reported that diluting quirk data with unrelated data reliably makes quirks harder to find, even at ratios as low as one to 0.1. Here, at a much heavier one-to-one ratio and with expression held constant, activation oracles showed interpretability going down in four cases and up in two. Mixing is not a dependable way to make a harder test.
Two more results worth keeping. Changing only the data generation pipeline, holding the quirk, method, and expression rate fixed, shifted interpretability substantially. And the SAE results were mostly swallowed by a high noise floor, which left little valid signal to compare at all.
The lottery
Here is the reading we take from this. Each organism is one draw from a wide distribution of construction choices that all produce the same visible behaviour at the same rate. A published auditing score is a score on one draw. If scores vary twentyfold across draws, and the ordering of draws changes from technique to technique, then a single organism cannot tell you how a technique will do on the next organism, let alone on a naturally arising behaviour in a production model.
The paper also shows that the diffing setting, where the auditor has the original model to compare against, does better than non-diffing whenever it clears the noise floor, and that the variation across training methods looks different in the two settings. That matters because the easiest published results tend to be diffing results, and a real audit of a deployed model will often lack a clean reference model to diff against.
A worked example of the failure mode. Suppose a team reports that their oracle identifies the implanted quirk in 90 percent of prompts on an organism built by synthetic document fine-tuning with no mixing. A reader will infer the oracle finds quirks. What the suite suggests is that the same oracle on the same quirk instilled through integrated DPO might land at a fraction of that, and nobody would know, because only one organism was built.
What we would change
The authors' recommendation is sensible and we would go a little further. Benchmarks should keep their quirk diversity and add construction diversity, reporting a technique's score as a range across methods rather than a number on one model. The 54 organisms and their training data are public, so this is no longer a matter of building things from scratch.
The one caveat we hold is scale. These are one billion parameter models and the authors say plainly that transfer to larger systems is unknown. It is possible that at frontier scale the construction method washes out and quirks look alike. It is also possible that the realism gap widens. Until someone repeats a slice of this suite at ten or seventy billion parameters, the safe assumption is that our auditing scores are partly measuring how we built the patient.
Sources
From the foundation