Turn-averaged SAEs: fewer features, and the ones you actually wanted
Anthropic's June update trains dictionaries on the residual stream averaged across a whole conversation turn. The volume of features to read drops from tokens times L0 to L0, and the features that surface describe behaviour rather than syntax. A small method with a large lesson for anyone auditing transcripts.
The problem with reading every token
A standard sparse autoencoder gives you a set of active features at every token position. If a dictionary has an L0 of around fifty and a transcript has two thousand tokens, that is a hundred thousand feature activations for one conversation, and a long agentic transcript runs into the millions. The June circuits update from Anthropic states the problem plainly. Even short transcripts produce thousands to millions of activations to interpret, and most of them describe things nobody auditing a model cares about, like the current token being a digit or the sentence being inside a code block.
This is the usability gap that has kept SAEs on the research side of the wall. The features exist, and many are interpretable, but the act of finding the ones that matter for a safety question means reading through an ocean of syntax. Every practical auditing pipeline we know of has ended up bolting some form of aggregation on top, and the aggregation is usually ad hoc.
Average first, then learn the dictionary
The update's answer is to move the aggregation before the dictionary. Take the residual stream at every token in a single conversational turn, either a Human turn or an Assistant turn, average it into one vector, and train the SAE on those turn vectors. The dictionary then never sees per-token structure. For a given turn it surfaces roughly L0 active features rather than the number of tokens multiplied by L0, which for a long turn is a reduction of three orders of magnitude.
They trained this on Qwen-2.5-7B-Instruct using the LMSYS-Chat-1M dataset of real user conversations. A detail we found reassuring is that the dictionary generalises to turns about 150 times longer than anything it saw during training. Averaging is a strong normaliser, so the input distribution does not drift much with length, and that is what you would hope for if the intended use is long agent transcripts.
The puzzle example
The worked example in the update is the one that convinced us the method is doing something different rather than just something cheaper. They took a prompt where the model gives a wrong answer to a numerical puzzle. A per-token SAE run on that turn lights up features about arithmetic and digits, which is correct and useless. You already knew the turn was about numbers. The turn-averaged dictionary surfaces features related to incorrect answers in number puzzles.
That is a behaviour-level description of the turn, and it is the kind of thing an auditor is asking for. The reason it appears is not mysterious. Averaging over the turn washes out anything that varies token by token and leaves what is constant across the turn, and the fact that the answer is wrong is a property of the whole response rather than of any one token. The dictionary is being trained on a representation that has already thrown away the syntax, so it has to spend its capacity on something else.
What is lost
The update is careful to frame this as a preliminary observation rather than a result, and we want to keep that framing. Averaging destroys position information, so anything about where in a turn something happened is gone. It also means the features cannot be attributed to specific tokens without going back to a per-token dictionary, which reintroduces the volume problem for the cases you drill into. We would expect the two dictionaries to be used together, coarse first and fine second, rather than one replacing the other.
There is also a question about what a turn-averaged feature is a feature of. A per-token SAE feature is a direction in a single residual vector and there is a decent theory of what that means. A turn-averaged feature is a direction in a mean of many residual vectors, and a mean can be high on a direction because one token was very high or because many tokens were slightly high. Those are different mechanisms with the same signature. Any causal story built on these features will need to check which one is happening.
The usability lesson
The lesson we take is that the unit of analysis should match the question. Auditing asks questions about turns, episodes and conversations, and for years we have been answering them with a tool whose native unit is the token, then aggregating afterwards. Training the dictionary at the unit you actually care about is a small change with a large payoff, and we suspect it generalises. Averaging over a tool call, over a reasoning block, or over an entire episode are obvious next things to try.
What we would want to see next is a direct comparison on an auditing task with ground truth. Take a set of transcripts with known problems, run both dictionaries, and measure how many features a reviewer has to read before they find the problem. That number, rather than reconstruction loss, is the metric this method is designed to move, and it is the one that would tell us whether the puzzle example is representative or lucky.
Sources
From the foundation