SAE probes in production: Rakuten's PII detector
Goodfire and Rakuten deployed probes on sparse autoencoder features to catch personal data in a multilingual agent platform. An applied look at what an interpretability method has to offer over a plain classifier to earn a place in a pipeline.
The deployment
Rakuten runs an agent platform in Japan serving over 44 million monthly active users, and it needs to catch personal names, addresses, phone numbers and email addresses before they leak into logs or downstream models. Goodfire's write-up describes the detector they built for this. A Llama 3.1 8B model runs as a sidecar. A sparse autoencoder trained on its layer 12 residual stream, with an expansion factor of 8 for 32,768 latents, turns each token's activation into a sparse feature vector. A classifier on top of those features flags the PII.
The SAE is a BatchTopK variant, trained on a mix of FineWeb and Japanese OpenWebText, with a sequence-wise activation at inference time. The classifier heads tested were logistic regression, random forests and XGBoost. Random forests did best on SAE features. XGBoost did best on raw activation probes. Two things about this setup are worth pausing on. Nothing is fine-tuned. The base model is frozen, the SAE is frozen, and only a small classifier is trained. And the training data was entirely synthetic, because nobody could be given real customer PII to build the detector with.
The result that justifies it
On real production data the SAE probe reached 96 percent F1, with recall around 92 percent on exact match and 96 percent on partial match. The same Llama 8B, used as a black-box judge by prompting it to find PII, scored 51 percent F1 on the same data. So reading the model's internals extracted a great deal more signal than asking the model. That is the cleanest number in the piece and the one we would lead with in any argument about why white-box access matters.
Against larger judges the story is about cost. Goodfire reports the probe at between 10 and 500 times cheaper than LLM-as-a-judge setups with comparable performance, with GPT-4o Mini and Claude Opus 4.1 as the comparisons. The probe runs on an 8B model and its marginal cost per additional probe is close to zero, since the SAE features are computed once and any number of classifiers can read them.
Where SAE features beat raw activations
The obvious alternative to an SAE probe is a probe on the raw residual stream, which skips the dictionary entirely. Goodfire tested both across three conditions. Synthetic data on the training distribution. Synthetic data with realistic formatting added. Real production data with its own formats and class imbalances. The SAE probe outperformed the activation probes and all the attention probes on all three. The gap opened widest on the shift from synthetic to real, where activation probes degraded sharply and SAE probes held.
The one exception was English-only data, where activation probes sometimes came close. On Japanese and mixed-language data the SAE advantage was decisive. The write-up leaves the why open. The plausible story is that a sparse dictionary trained on both languages gives the classifier features that mean the same thing in either, whereas a linear probe on dense activations picks up whatever direction happened to separate the synthetic examples. With a training set that is entirely synthetic, generalisation under shift is the whole problem, and this is where the dictionary earned its keep.
What fine-tuning would have given up
Goodfire is candid that a fine-tuned model matched the SAE probe on exact-match performance. So the case for the probe rests elsewhere. Probes are lightweight, adaptable and need very little data. A new PII class means training a new classifier on the same frozen features, with a few examples, in minutes. A fine-tuned detector means a new training run and a new model to serve. When the platform is multilingual and the categories of sensitive data shift with regulation, that difference in iteration cost matters more than a point of F1.
There is a second advantage the write-up does not dwell on but which we think is real. An SAE probe is inspectable. If it starts flagging something it should not, you can look at which features carry the weight and often read what they represent. A fine-tuned classifier gives you a score and nothing else. In a pipeline that handles customer data under privacy law, being able to say why a string was flagged is worth something on its own.
What this does and does not show
This is one deployment, on one task, reported by the company that sells the tooling, and the baselines were chosen by them. PII detection is also a task where the relevant concepts are concrete and well represented in web text, which is exactly where dictionary features should be easiest to find. We would not extrapolate to detecting deception or intent, where the features are harder to name and harder to trust.
What it does show is that SAE features can survive contact with production traffic and beat the practical alternatives on cost and shift robustness. That is more than the method had demonstrated before. The experiment we would want next is a replication on a second task with a second SAE, trained by someone other than the vendor, and a report of how many features the winning random forest actually leaned on. If the answer is a few dozen interpretable ones, the case for this approach gets much stronger.
Sources
From the foundation