The result

Rulin Shao, Shuyue Stella Li and colleagues at Washington and AI2 ran reinforcement learning with verifiable rewards on Qwen2.5-Math-7B and varied only the reward. Ground truth rewards gained 29.1 points on MATH-500. Majority vote gained 27.1. Rewarding the presence of a boxed answer, with no check on whether it was right, gained 13.8. Deliberately incorrect labels gained 24.1. Random rewards, a coin flip with no relation to the answer, gained 21.4.

Read that last one again. A reward signal carrying no information about correctness produced roughly three quarters of the gain of a correct one. If reinforcement learning is supposed to be moving the policy toward behaviour the reward prefers, and the reward prefers nothing in particular, then something other than the reward is doing the work.

The same recipe on Llama3.1-8B, Llama3.2-3B and OLMo2-7B produced nearly nothing. Spurious rewards helped Qwen models and only Qwen models, which is what turns a strange result into a diagnosable one.

What Qwen was already doing

The authors looked at what the base model does before any RL. Qwen2.5-Math-7B writes Python code to help itself reason in about 65 percent of its responses, despite having no way to execute it. That behaviour is worth a lot: responses that use code are 60.9 percent accurate, responses that do not are 28.0 percent accurate.

After RLVR, code reasoning appears in around 90 percent of responses under every reward condition, including the spurious ones. So the training raises the frequency of a behaviour the model already had and already did better with, without teaching it anything new. Any reward that nudges the distribution toward the model's own high-frequency, high-accuracy mode will look like it is teaching mathematics.

That also explains the model dependence. Llama and OLMo do not have this pre-existing code-reasoning habit at anything like the same rate, so there is no latent behaviour for a noise signal to amplify.

How a random reward becomes a gradient

The mechanism the paper proposes is the clipping in GRPO. In a PPO-style objective, the ratio between the new and old policy is clipped, and the clipping does not act symmetrically across tokens. Tokens the model already assigns high probability are clipped less often and accumulate positive updates. Low-probability tokens hit the clip boundary more often and get penalized. The net effect concentrates the policy on what it already does, regardless of what the reward says.

The test is the one you would want. Disable clipping and rerun. Across their conditions, random rewards then produce no reliable improvement. That is a clean ablation and it turns a story about the reward into a story about the optimizer.

The part that stays open

The authors are direct about what they have not shown. Their analysis centres on code reasoning, and other latent behaviours could be contributing without being measured. They note that Qwen2.5-Math-7B is unusually sensitive to prompts, including prompts irrelevant to the task, which is its own warning about how stable any of these numbers are. And the mechanism behind spurious reward gains is not fully characterized even with the clipping result.

The question everyone asks next is contamination. If the base model has seen MATH-500 or material very close to it, then RLVR on any signal is eliciting memorized content and the whole comparison changes. The paper does not settle this, and the honest position in June 2025 is that the elicitation story and the contamination story both fit the evidence, and they are not mutually exclusive. A model can be amplifying a genuine reasoning habit that happens to have been trained on the test distribution.

Either way the consequence for practice is the same. An RLVR result reported on a single model family is uninterpretable. If your reward design gains points on Qwen and nobody has run it on Llama, OLMo or anything else, you do not yet know whether you improved the method or found another way to amplify what Qwen already does.

What we would ask of the next RLVR paper

Report the random-reward baseline. It costs one extra run and it tells the reader how much of the reported gain requires the reward to be correct. A method that beats random rewards by two points on one model family is a different claim from one that beats them by twenty.

Report the base model's behaviour distribution before and after, the way this paper reports code-reasoning frequency. Most of the RLVR gains we have seen in the last year could be described as sharpening rather than learning, and nobody can tell which without that measurement. And run at least two model families. This paper is a replication study, and the reason it found something is that it did the thing most of the field skipped.

Sources

  1. Shao et al., Spurious Rewards: Rethinking Training Signals in RLVR (arXiv 2506.10947)
  2. Lab review: Spurious Rewards, rethinking training signals in RLVR