s1: a thousand examples and the word 'Wait'
Muennighoff and colleagues fine-tuned Qwen2.5-32B-Instruct on 1,000 curated reasoning traces in 26 minutes on 16 H100s and beat o1-preview on competition math. Notes on how little data it took, what the 'Wait' trick reveals about where reasoning ability already lives, and where the method runs out.
The result
The s1 paper from Stanford and the University of Washington makes a claim that sounds like a typo. Take Qwen2.5-32B-Instruct, fine-tune it on 1,000 questions with reasoning traces, and the resulting model scores 56.7 percent on AIME24 against 44.6 for o1-preview and 26.7 for the base model. On MATH500 it gets 93.0 against o1-preview's 85.5. On GPQA Diamond it gets 59.6, which is well above the base model's 49.0 but well below o1-preview's 73.3, so the story is math and not everything.
The training run is the part that made us reread the paper. Supervised fine-tuning for five epochs, batch size 16, learning rate 1e-5, on 16 H100s, taking 26 minutes. The authors, who include Niklas Muennighoff, Percy Liang, Tatsunori Hashimoto, Emmanuel Candès and Li Fei-Fei among others, released the model, the data and the code. Anyone with two nodes can rerun it before lunch.
How the thousand were chosen
The dataset is where the work is. The team started with 59,029 questions from 16 sources, the largest being NuminaMATH at 30,660, then MATH at 11,999, OlympicArena at 4,250, OmniMath at 4,238 and AGIEval at 2,385. They added two of their own: s1-prob from Stanford statistics qualifying exams and s1-teasers from interview brain-teasers. Each question got a reasoning trace generated by Gemini.
Three filters took the pool down. A quality pass removed formatting problems and left 51,581. A difficulty pass ran Qwen2.5-7B and Qwen2.5-32B on each question and kept only those both got wrong or that produced long traces, leaving 24,496. A diversity pass classified questions into domains with Claude 3.5 Sonnet using the Mathematics Subject Classification scheme and sampled across 50 domains to reach 1,000.
The ablations justify each step. A random 1,000 from the pool gets 36.7 on AIME24. A diversity-only selection gets 26.7. Selecting the 1,000 longest traces gets 33.3. The combined s1K selection gets 50.0 before any test-time trick. And training on all 59,000 examples gets 53.3, at a cost of 394 H100 hours against 7 for s1K. That last comparison is the one to remember: 59 times the training data bought three points.
Budget forcing and the word Wait
The second ingredient is a decoding intervention the authors call budget forcing. To cap thinking, they insert the end-of-thinking delimiter and the model moves to its answer. To extend thinking, they suppress the delimiter when the model tries to stop and append the string Wait, which makes the model reread its own reasoning and often catch a mistake. Applying Wait twice takes s1-32B from 50.0 to 56.7 on AIME24.
They tried other strings. Wait worked best, which we find both funny and revealing. The model was never trained to respond to that token in that position. It works because the base model, and the thousand traces layered on it, already contain the habit of double-checking when prompted, and a single word is enough to trigger it. That is the strongest evidence in the paper for the claim that the reasoning ability was mostly in Qwen2.5-32B before anyone fine-tuned it, and the thousand examples taught it a format rather than a skill.
The authors also compared sequential scaling against parallel scaling. Majority voting over 64 samples from the base model does worse than s1-32B with budget forcing. Their argument is that sequential computation can build on intermediate results while parallel samples cannot, and the numbers support it at this scale.
Where it stops working
The paper is unusually direct about limits. Budget forcing flattens out. After four to six applications of Wait, accuracy stops improving and the model sometimes falls into repetitive loops rather than doing more reasoning. The context window of the base model caps how far the trick can go at all. So this is a way to expose test-time scaling, not a way to get unbounded returns from it.
The GPQA gap is the other caveat. On competition math the fine-tune plus budget forcing beat o1-preview. On graduate-level science questions it did not come close. The thousand examples are heavy on mathematics by construction, and whatever o1-preview was trained on covers more ground. A thousand examples can surface an ability that pretraining already built. They cannot build one that is missing.
What we take from it
This is a replication story more than a method paper. The reasoning models released in late 2024 came with the implication that the capability required a large reinforcement learning effort. s1 shows that a substantial fraction of the visible gain on math benchmarks can be recovered with supervised fine-tuning on a carefully chosen sample, in under half an hour, from a base model anyone can download. The GitHub repository already lists an s1.1 release with traces regenerated from DeepSeek-R1 instead of Gemini, and recommends it over the original.
The experiment we want someone to run is the difficulty filter alone at larger sizes. If 1,000 examples give 50 and 59,000 give 53, there is a curve between them, and its shape would tell us whether the thousand is a floor set by format learning or a sweet spot set by something else. The second thing to try is Wait on models that were never fine-tuned at all, with the delimiter trick, to measure how much of the effect is the data and how much is the decoding. The paper gives the tools. The rest is a week of GPU time.
Sources
From the foundation