o1 and the new axis: scaling test-time compute
OpenAI says o1 gets better the longer it thinks, and that it was trained with reinforcement learning on its own chain of thought. Here is what that claim means and what remains hidden.
The claim
OpenAI released o1-preview and o1-mini on September 12 with a short technical post and two plots. One plot shows accuracy on AIME rising with train-time reinforcement learning compute. The other shows accuracy rising with test-time compute, meaning the number of tokens the model spends thinking before it answers. Both axes are log scale and both curves look smooth. That second plot is the reason people are calling this a new scaling law.
The description of the training method is one sentence. OpenAI says its large-scale reinforcement learning algorithm teaches the model how to think productively using its chain of thought in a highly data-efficient training process. The post says the model learns to recognize and correct its mistakes, to break hard steps into simpler ones, and to try a different approach when the current one is not working. On the benchmark side, OpenAI reports that o1 places in the 89th percentile on Codeforces, ranks among the top 500 US students on the AIME qualifier, and exceeds PhD-level human accuracy on GPQA.
What RL on chain of thought is training
Ordinary chain of thought prompting asks a pretrained model to write its reasoning out and hopes the reasoning is good. The pretraining loss never rewarded the reasoning for being correct, only for looking like text. RL on chain of thought closes that loop. The model produces a long reasoning trace, the final answer is checked against ground truth, and the trace that led to a correct answer gets reinforced. The reasoning itself becomes the thing under optimization.
This explains the pattern in OpenAI's results better than any architecture story. Gains are large on math, code and standardized tests, where an answer can be checked mechanically, and the post reports little change on tasks like creative writing, where there is no clean reward. The technical primer on LessWrong makes the same point. Everything here is a post-training change on top of an ordinary language model, and the model gets good at exactly the things a verifier can score.
The primer lists four families of method that could produce this, and we think it is the right frame for reading the release. One is to sample many attempts, keep the ones a verifier accepts, and finetune on those. One is to train a process reward model that scores partial traces and optimize against it. One is to use a verifier to guide search, in the style of beam search or Monte Carlo tree search, at training time. One is to build traces that contain a mistake and its correction and train on those. OpenAI has not said which of these it uses, or in what combination. The behaviours it describes, backtracking and self-correction, are consistent with several of them.
What is hidden and why
The chain of thought is not shown to users. OpenAI gives two reasons. The first is that it wants the model to reason freely about policy compliance without that reasoning being visible, and in its own words it cannot train any policy compliance or user preferences onto the chain of thought and does not want to make an unaligned chain of thought directly visible. The second, which Simon Willison draws out, is competitive. Showing the traces would let other labs train on them.
The traces still cost money. Reasoning tokens are billed as output tokens even though the API does not return them, and OpenAI suggests budgeting around 25,000 of them for hard prompts. Output limits went up to match, 32,768 tokens for o1-preview and 65,536 for o1-mini against 16,384 for GPT-4o. So the test-time scaling curve is also a pricing curve. More accuracy costs more tokens, and the customer pays for tokens they cannot inspect.
For a researcher this is frustrating in a specific way. The test-time plot is the interesting scientific result and we cannot check it. We do not know the token counts on the x axis, whether the curve flattens, or whether the same curve appears on tasks outside the ones OpenAI chose. The right response is to reproduce it, and we expect that to happen quickly because the recipe is not exotic.
Why replications will come fast
Every ingredient is available in the open. Strong base models exist. Verifiable math and code datasets exist. RL infrastructure for language models exists and has been used for RLHF for two years. The thing OpenAI did that others had not is commit to a long RL run with a correctness reward and let the model write as much as it wants. That is an engineering and budget decision more than a secret.
Our prediction is that within a few months there will be at least one open model with a visible chain of thought that shows the same test-time curve on AIME, and that some of the replications will use far less than OpenAI's compute because they will start from a strong open base and a small, curated set of reasoning traces. The interesting question is whether the curve is a property of the RL or a property of long-form reasoning that a small amount of finetuning can bring out.
What we would want to know
We would like the x axis of the test-time plot published in tokens. We would like the same plot on a task with a noisy reward, to see whether the curve survives outside math. And we would like an ablation that holds the number of reasoning tokens fixed and varies the amount of RL, because that would separate the two scaling axes OpenAI has drawn on the same page. None of that requires OpenAI. It requires an open replication, and the labs racing to build one will answer these questions faster than the original authors will.
Sources
From the foundation