Features as rewards: using SAE features as an RL signal against hallucination
Goodfire trained Gemma-3-12B-IT with reinforcement learning where the reward came from probes on the model's own activations, and cut hallucinations on a held-out set by 58 percent. Notes on interpretability crossing from diagnosis into training, and on why the Goodhart question is still open.
Reading the activations instead of the output
Goodfire published a post-training method on February 11 that uses interpretability tooling as the reward, and the headline number is a 58 percent reduction in hallucination rate on a held-out test set. The method is called Reinforcement Learning from Feature Rewards, or RLFR. The idea is that the model already carries signals in its activations about whether a claim it is producing is grounded, and those signals can be read by lightweight probes and turned into a training signal rather than only a monitor.
We want to be precise about what the reward is, because the name invites a misreading. The reward comes from probes trained on internal representations rather than from a single sparse autoencoder feature that someone picked out of a dictionary. The distinction matters for the Goodhart discussion later. A probe is a small classifier on activations. It can be trained to a target label. A hand-picked SAE feature is fixed and its meaning is whatever the dictionary learned.
The four-stage pipeline
The setup has four stages. A token-level localisation probe identifies candidate entity spans in the model's output. A span-level classification probe flags which of those spans are hallucinated. The policy then generates inline corrections or retractions for flagged spans. Reward probes grade the quality of those interventions and that grade is the RL reward.
The probes were trained on a curated dataset with labels produced by Gemini 2.5 plus web search. The model being trained was Gemma-3-12B-IT, and the RL prompts came from LongFact++, a set of 20,000 knowledge-intensive prompts across eight domains. The whole training run cost about 2,500 dollars for roughly 360 optimiser steps.
That cost is the argument for the method. The authors estimate that using Gemini 2.5 Pro with web search as the reward signal would have bought only three to five steps for the same money, which they put at about 90 times higher cost per intervention. Probes are cheap to run once trained, and the training of the probes happens once, up front.
What the numbers break down into
The 58 percent figure uses best-of-32 sampling. Without best-of-n the reduction is 31 percent. The authors decompose the effect into a 10 percent reduction from the policy itself, a 35 percent in-context reduction, and a 12.5 percent direct reduction. They report no degradation on standard benchmarks, and they show that intervention success rises as more samples are drawn, which is a test-time scaling curve for a safety property rather than for accuracy.
The held-out set is 999 prompts from LongFact++. Since the training prompts and the evaluation prompts come from the same distribution, this is a within-distribution result. What it says about hallucination on a conversational assistant workload, or on code, is not yet measured.
Where the Goodhart question actually sits
The obvious worry with optimising against a signal you can read from activations is that the policy will learn to change its representations rather than its behaviour, and the probe will go blind. The authors name this directly. Their design choice is that the reward probes are run on a frozen copy of the model instead of on the student's activations. No gradient flows through a probe, so the student is rewarded for producing tokens that a fixed reader scores well, and it cannot reach the reader's activations to fool it directly.
That is a real safeguard against one failure. It leaves open whether the frozen reader can be gamed through the tokens alone. A policy could in principle learn output patterns that make the frozen model's activations look grounded without being grounded. The authors' evidence against this is that probes trained on base model activations stayed effective on the trained policy's activations, which they read as the representations remaining stable while behaviour changed. Their phrase is that changing behaviour turned out to be far easier than changing representations, at least at this level of optimisation.
We read that as a hopeful result with a narrow scope. 360 steps is a short run. The honest claim is that under this much pressure, on this model, with the probes frozen, the signal held. Whether it holds at ten times the optimisation, or when the probe is the only reward for a long horizon, is the experiment we would want to see next.
Diagnosis to training
For a few years the pitch for interpretability was mostly about audit. You train the model, then you look inside to check what it learned. RLFR moves the reading of internals into the training loop, which is a different relationship. A monitor that is wrong sometimes costs you a false alarm. A reward that is wrong sometimes teaches the model something.
The dependence on probe quality is the limitation the authors list, and it is the one we would put first. The probes here are trained against labels from a large model with search access, so the ceiling on what the reward can know is the ceiling of that labelling process. If we were reproducing this, we would start by measuring how much of the 58 percent survives when the probe labels are noisier, and by running the frozen reader on an open model where anyone can check whether the representations really did stay put.
Sources
From the foundation