The spurious rewards paradox gets a mechanism
A January paper traces why RLVR with wrong rewards still improves Qwen 2.5 to an anchor-adapter circuit in the middle and late layers that retrieves memorised answers. Notes on the perplexity paradox, the circuit, and what it means for contamination detection.
The puzzle from last summer
In mid 2025 a widely discussed result showed that reinforcement learning with verifiable rewards could improve Qwen 2.5 on maths benchmarks even when the reward signal was random or systematically wrong. That was hard to square with any story in which RL teaches reasoning. If the reward carries no information about correctness, whatever improved cannot have been learned from the reward. The leading explanation was that the training was surfacing something the model already had, and the obvious candidate was memorised test data.
A paper submitted on January 16 by Lecheng Yan, Ruizhe Li, Guanhua Chen and five coauthors takes that explanation and tries to make it mechanistic. Their title calls the phenomenon the spurious rewards paradox, and the contribution is to locate where in the network the shortcut lives and to show that it can be turned up and down by hand.
The perplexity paradox
The first observation is a training signature. After spurious RLVR, the model's perplexity on the answer tokens of the test set goes down. At the same time, the coherence of its handling of the prompt degrades. A model that had learned better reasoning would be expected to engage with the problem statement more carefully, not less. A model that is retrieving stored answers can afford to skim the prompt, because the prompt is only a key for the lookup.
The authors call this the perplexity paradox and use it as the primary evidence that the gains are memorisation rather than reasoning. It is a cheap diagnostic. Any lab running RLVR on a model with an unknown pretraining corpus can measure answer-token perplexity on the evaluation set before and after training. If it drops sharply while prompt-side behaviour gets sloppier, the improvement is suspect. We would like to see this become a routine check in RL reports.
The anchor-adapter circuit
The mechanistic work uses path patching, the logit lens, Jensen-Shannon divergence between layer distributions, and a neural differential equation view of the residual stream to trace how the shortcut is computed. The picture that emerges has two parts. A functional anchor sits in the middle layers, around layers 18 to 20 in the models studied, and acts as the trigger that initiates retrieval of a memorised solution. A set of structural adapters in layers 21 onward then reshapes the representation so that the retrieved content is expressed through the output pathway.
The key test is intervention. The authors identify specific MLP keys within this circuit and show that scaling them amplifies or suppresses the contamination-driven performance. Turn the keys up and the model gets better on contaminated items without any change to its general ability. Turn them down and the spurious-reward gains disappear. That is the kind of result that separates a description from a mechanism. If you can dial the behaviour with a handful of parameters, you have found the parameters that matter.
From puzzle to detector
What we find most useful is the reframing. A year ago the spurious rewards result was an embarrassment for the RLVR story and an argument about whether Qwen was special. Now it reads as a contamination detector that happens to use RL as the probe. Train briefly with a random reward, watch the answer-token perplexity, and look at whether activity in the anchor layers rises on the evaluation set. If it does, the evaluation set is in the model.
The practical version is even simpler. Suppress the identified MLP keys and re-evaluate. Whatever score survives suppression is the part of the benchmark result that does not depend on retrieval. That is a more honest number than the headline score, and it can be computed without knowing what the pretraining corpus contained. Contamination has been hard to detect precisely because nobody outside the lab can inspect the corpus, so an internal, weight-level probe is the right shape of tool.
What we are not yet sure of
The circuit is described for the Qwen 2.5 family, and the layer indices are specific to those models. Whether other families that also show spurious-reward gains have a similar two-stage structure, or whether the anchor lands in a different place, is not settled by this paper. We would also want to know how the circuit behaves on items that are near duplicates of training data rather than exact copies, because that is where most real contamination lives.
The reproduction we want to run is on a model with a fully open pretraining corpus, where contamination can be established by string matching rather than inferred. If the anchor-adapter signature fires exactly on the items we can prove were seen, and stays quiet on the rest, then the method is calibrated and can be pointed at closed models with some confidence. The code is public, and this seems like a week of work for a small team. We would be glad to compare notes with anyone who tries it.
Sources
From the foundation