DeepSeek-R1: reasoning from pure reinforcement learning
R1-Zero learned to reason with reinforcement learning and no supervised examples, and R1 matches o1 on the benchmarks that matter. The distilled small models are the part we keep coming back to.
What was released
DeepSeek published the R1 paper on January 22 along with open weights for two large models and six small ones. R1 scores 79.8% pass@1 on AIME 2024 against 79.2% for OpenAI's o1-1217, 97.3% on MATH-500 against 96.4%, and holds a Codeforces rating of 2029. Those numbers are the headline. The method is the story, and the method is mostly reinforcement learning with a very simple reward.
The paper describes two models. R1-Zero is trained by RL directly on DeepSeek-V3-Base with no supervised finetuning at all. R1 is the production version, which adds a small cold-start dataset and two more stages on top. The gap between them tells you what each piece of the recipe buys.
What R1-Zero shows about where reasoning comes from
R1-Zero uses GRPO, a policy gradient method that scores a group of sampled responses to the same prompt against each other rather than against a learned value function. The reward has two parts. An accuracy reward checks whether the final answer is right, using rule-based verification for math and test execution for code. A format reward checks that the model put its reasoning between the expected tags. The authors say they avoided neural reward models because of reward hacking.
That is the whole signal. No human-written reasoning traces, no process supervision, no examples of what good thinking looks like. From that, R1-Zero reaches 71.0% pass@1 on AIME 2024. Along the way the response length grows on its own, and the paper describes what it calls an aha moment, where the model starts to reevaluate its initial approach and allocate more thinking time to a problem without being told to. Self-verification and reflection appear as behaviours the reward found useful, not as behaviours anyone wrote down.
We find this the most important result in the paper. It says that a strong base model already contains the pieces of reasoning, and that a correctness signal is enough to assemble them. The pretraining put the capability there. The RL taught the model when to use it. That is a different picture from the one where reasoning has to be taught by example, and it is a much cheaper picture.
What the extra stages buy
R1-Zero has problems the paper is candid about, including poor readability and language mixing in its traces. R1 fixes those with a pipeline. First a few thousand long chain-of-thought examples as a cold start, so the model begins RL with a readable format. Then reasoning-oriented RL as before. Then rejection sampling from that checkpoint to build a supervised set of about 600k reasoning samples and 200k non-reasoning samples, and a finetune on it. Then a final RL pass covering all scenarios, including helpfulness and harmlessness.
The AIME gain from R1-Zero to R1 is 71.0% to 79.8%. That is real but it is the smaller step. The bigger step was the one from the base model to R1-Zero, and it was taken by RL alone. The cold start and the SFT stages look to us like they are mostly about making the model usable, with a modest accuracy bonus on top.
Epoch AI's Ege Erdil has tried to estimate what the RL stage cost. Working backwards from the paper's stated 8000 gradient steps and reasonable guesses for batch size and rollouts, he lands at roughly 6.1e23 FLOP, or around one million dollars at pretraining-like utilization, against his estimate of about 5.3 million dollars for the V3 pretraining itself. He is careful to say batch size and rollout count are not stated, so the number could be off by a factor of two or more. Even so, the shape is clear. The reasoning stage cost a fraction of the base model.
Why the distilled models matter as much
Alongside R1, DeepSeek released six dense models at 1.5B, 7B, 8B, 14B, 32B and 70B parameters, built by finetuning Qwen and Llama checkpoints on the 800k samples from R1. The 32B distilled model scores 72.6% on AIME 2024. That is above R1-Zero, and it is a model most research groups can run on a single node.
The paper includes a comparison that we think will shape a lot of work this year. They tried running the RL recipe directly on a small model and found that distilling from R1 substantially outperforms it. Reasoning can be discovered by RL at large scale and then transferred by supervised finetuning to small scale, but the small model cannot easily discover it on its own. For anyone without a frontier budget that is the practical lesson, and DeepSeek made it usable by shipping the checkpoints.
It also reframes what open weights are for. The big model is the demonstration. The small ones are the tool. A 7B model that reasons at a level that required a proprietary API a few months ago changes what a university lab or a non-profit like ours can study directly, including the failure modes.
What we would want checked
Three things. First, the aha moment is described with a single example in the paper and we would like to see it quantified, for example by measuring how often reflective phrases appear as a function of training step and whether they correlate with correctness. Second, the claim that RL on small models underperforms distillation is based on one comparison, and the RL run on the small model may simply have been under-tuned. Third, the accuracy reward is only available where an answer can be checked, and the paper's strongest numbers are all in those domains. How much of the reasoning transfers to tasks with no verifier is the open question, and the released weights make it answerable.
A note added later. The R1 paper was published in Nature in September 2025 after peer review, the first widely used frontier model to go through that process. The peer-reviewed version disclosed that the R1 RL stage cost about 294,000 dollars on 512 H800 GPUs over roughly 80 hours, which lands inside the range Erdil estimated in January. It is a rare case of a training cost claim being checkable against an independent estimate made before the disclosure.
Sources
From the foundation