Does RL create reasoning or just find it? The pass@k argument
Yue and colleagues at Tsinghua sampled base and RL-trained models hundreds of times per problem and found the base models solve more problems at large k. Reading notes on what that means for what post-training buys.
The question and the metric
The story of the year so far is that reinforcement learning with verifiable rewards teaches models to reason. DeepSeek-R1 and the models that followed get their gains from RL on math and code with a binary correctness signal, and the natural reading is that the model discovers strategies it did not have, the way AlphaGo discovered moves. A paper from Yang Yue, Gao Huang and colleagues at Tsinghua's LeapLab tests that reading directly and finds it mostly wrong for current methods.
Their instrument is pass@k. Sample k answers to a problem, count the problem solved if any one of them is correct, average over the dataset. At k equal to one this is ordinary accuracy. At k in the hundreds or thousands it measures something else: the set of problems the model can solve at all, given enough tries. The authors call this the reasoning capacity boundary, and the whole paper is a comparison of that boundary before and after RL.
What the curves show
On mathematics they use Qwen2.5 base models at 7, 14 and 32 billion parameters and Llama 3.1 8B, each paired with its zero-RL counterpart trained by SimpleRLZoo with GRPO, evaluated on GSM8K, MATH500, Minerva, Olympiad, AIME24 and AMC23. At small k the RL model wins everywhere, which is the familiar result. As k grows the curves cross, and at k of 128 or 1,024 the base model is ahead on every benchmark and every model family. On Minerva with the 32B model the base model is about 9 percent ahead at k equal to 128, which they read as 9 percent more problems in the validation set that the base model can solve and the RL model cannot.
The same crossing appears in code, comparing Qwen2.5-7B-Instruct-1M with its Code-R1 descendant on LiveCodeBench, HumanEval+ and MBPP+, and DeepSeek-R1-Distill-Qwen-14B with DeepCoder-14B. It appears in visual reasoning too, with Qwen2.5-VL-7B on MathVista and MathVision. Code is the cleanest case, because a solution has to pass unit tests and cannot be reached by lucky guessing. For math, where a wrong chain of thought can land on a right number, they manually checked the hardest problems: on the GSM8K questions with average accuracy under 5 percent, the base model got 25 right, 24 of them with at least one correct chain of thought.
Narrowing rather than expanding
The accuracy histograms explain the crossing. After RL the distribution of per-problem accuracy shifts toward one, which is the sampling efficiency gain, and simultaneously piles up at zero, which is problems the model has stopped being able to solve. Table 2 makes it concrete on AIME24 at k equal to 1,024: 63.3 percent of problems are solved by both base and RL model, 23.3 percent by neither, 13.3 percent by base only, and 0 percent by RL only. On MATH500 the RL-only cell is 1 percent.
The perplexity analysis closes the loop. They score RL-model responses under the base model and find they sit in the low-perplexity part of the base model's own distribution, meaning those responses were already likely under the base model. Responses from OpenAI o1, by contrast, have high perplexity under the base model, which is what genuinely new reasoning would look like. Their summary is that current RLVR sharpens the distribution within the base model's prior rather than pushing beyond it. Distillation does push beyond it: DeepSeek-R1-Distill-Qwen-7B's pass@k curve sits above its base model at every k.
Which algorithm does not matter much
A useful side result is the comparison of RL algorithms. They reimplemented PPO, GRPO, Reinforce++, RLOO, ReMax and DAPO in the same framework and measured what they call the sampling efficiency gap, the distance between an RL model's pass@1 and the base model's pass@256. On an in-domain Omni-MATH split the gap ranges from 43.9 for GRPO to 42.6 for RLOO. The spread between algorithms is a point or two. The gap to the base model's ceiling is above 40 for all of them. If the ceiling is set by the base model, then arguments about which policy gradient variant to use are arguments about a small fraction of the available gain.
What we believe and what we do not
The result we find hard to argue with is the narrowing. RL as currently practised makes a model better at the problems it could already solve and worse at a slice of the ones it could barely solve, and the effect is visible across three domains and several model families. That is a real cost that pass@1 evaluations hide, and anyone doing RL post-training should be plotting pass@k curves against their base model.
The stronger claim, that RL cannot create new reasoning, is a claim about the training regimes tested and we would hold it more loosely. The RL runs here are short and use single-turn, outcome-only rewards on a fixed problem set. The authors say as much, and their closing list of what might change the picture, better exploration, continued data scaling, process rewards and multi-turn interaction, is a research agenda rather than a concession. It is also possible that large k on a base model with a 16,000-token budget is doing a lot of search that the RL model has learned to do internally, so the comparison is between a cheap sampler and an expensive one rather than between two capacities. The experiment we want is the same pass@k comparison at ten or a hundred times the RL compute, on the same base model, with the crossing point reported. If it moves out past k equal to 1,024, the story changes. If it does not, the field has been buying sampling efficiency and calling it reasoning.
Sources
From the foundation