The setup

The paper from Adam Karvonen, Sam Marks, Owain Evans and a large group of collaborators, posted December 17, trains what they call activation oracles. Take a vector from the residual stream of a target model at some layer and position. Inject it into the oracle's own residual stream at layer one, in place of a reserved token, with the prompt stating which layer of the target it came from. Then ask the oracle a question in plain English and read its answer. The oracle is an ordinary language model fine-tuned to treat activations as one more input modality alongside text.

This builds on LatentQA, which did the same thing for narrow tasks. The contribution here is generality. The authors train on three kinds of data. System prompt QA takes transcripts where the target was given a system prompt such as act like a pirate, extracts activations from the conversation, and trains the oracle to answer questions about the target's disposition. Seven classification datasets are converted to yes or no questions. And a large self-supervised set asks the oracle to predict the text before or after the activation was taken. None of that training involves the models or secrets used in evaluation.

What it got right

Evaluation reuses four auditing tasks from prior work. Three are secret elicitation, including the Taboo game where a model has been fine-tuned to hint at a hidden word without saying it, and a secret side constraint task that also allows a black-box prompting comparison. The fourth is detecting emergent misalignment in a fine-tuned model, treated as a model diffing problem. Against existing white-box techniques, oracles matched or exceeded the best prior method on three of the four. The authors write that they achieve state of the art on three out of four auditing tasks, and the arXiv abstract phrases the weaker claim that oracles match or exceed white-box baselines on all four.

Two other findings seem more important to us than the leaderboard. Performance rises consistently with both the amount and the diversity of training data, across Qwen3-8B, Gemma-2 9B, Llama-3.3 70B and Claude Haiku 3.5 as targets. And training on more task types improves performance on the held-out auditing tasks, which is the pattern you would expect if the oracle is learning something general about how these models encode information rather than memorising a task.

An alternative to dictionaries

The obvious comparison is with sparse autoencoders. A dictionary gives you a fixed set of features with labels attached once, and reading a model means projecting activations onto that basis and looking up the labels. An oracle gives you a question answering interface over the raw vector. You lose the exhaustive inventory, since the oracle only tells you about what you ask, and the authors list that as a limitation. You gain the ability to ask about properties nobody anticipated when the dictionary was trained, and you avoid the reconstruction loss and the dead and split features that made dictionaries frustrating in practice.

The cost profile is different too. An oracle answer takes a forward pass of a model the size of the target or larger, and multiple questions mean multiple passes. The authors call inference expensive. Dictionary lookups are cheap once the dictionary exists, but the dictionary cost a large training run per layer. For auditing a handful of fine-tunes, where you want to ask a few open questions, the oracle economics look better. For monitoring every token in production, they do not.

The paper is also clear that oracles cannot answer mechanistic questions. They can tell you that an activation carries the concept of a hidden word or an unsafe disposition. They cannot tell you how the target computes it, which circuit is involved, or what would change it. That is a description tool, not an explanation tool, and it should be used as one.

The obvious worry

The worry every reader will have is that a language model asked what another model is thinking will simply answer from its own understanding. The authors raise it themselves. They note that an oracle trained to predict user demographics might learn to assemble that prediction on its own, whether or not the target model represents it, and that oracles might confabulate plausible guesses. Their phrase is that explanations may reflect the oracle's understanding rather than the target's.

We do not think this is disqualifying. It does move the burden. Every oracle claim needs a control where the activation is swapped for one from a matched context that lacks the property, and the oracle should answer differently. The secret elicitation tasks have this structure built in, since the secret was planted and the oracle had no other route to it, and that is why those results are convincing. Open-ended questions about a model's goals or beliefs do not have a ground truth, and an oracle's fluent answer to them is exactly the kind of output that should be treated as a hypothesis to test, not a reading to report.

What to try

The experiment we would run first is a calibration study. Ask the oracle a fixed set of questions about activations where the answer is known by construction, vary how strongly the property is present, and plot oracle confidence against the truth. If the curve is monotone the tool is usable with thresholds. If the oracle is equally fluent at every strength, its answers are prose and should be weighted accordingly.

The second is adversarial. Fine-tune a target to hold a secret while also being trained against an oracle's ability to read it, and see how much capability the target has to sacrifice. Dictionaries have never faced that test either. Whichever tool survives it is the one we would want in an audit.

Sources

  1. Karvonen et al.: Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers (arXiv 2512.15674)
  2. Anthropic Alignment Science blog: Activation Oracles