What shipped

On 20 January DeepSeek released R1, a reasoning model with open weights and a technical report describing how it was trained. A week later, on 27 January, the Nasdaq Composite lost more than a trillion dollars in a single day, with Nvidia among the names hit hardest. We have spent the week between those two dates reading the report and running the distilled models on our own hardware, and this note is an attempt to separate what changed from what the market thought changed.

The technical claim in the report is the important one. DeepSeek trained a variant, R1-Zero, using "pure reinforcement learning (RL), obviating the need for human-labeled reasoning trajectories," and found that the model developed "advanced reasoning patterns, such as self-reflection, verification, and dynamic strategy adaptation" on its own. The full R1 adds a small amount of supervised data on top of that. The report also states that the reasoning behaviour of the large model "can be systematically harnessed to guide and enhance the reasoning capabilities of smaller models," and DeepSeek released a family of distilled models alongside R1 to prove it.

Until this month, the recipe for o1-class reasoning was a rumour. Now there is a paper, a set of weights, and a price list. Mercer, Spillard and Martin, in an early analysis posted on 4 February, describe a model "developed at a fraction of the cost yet remains competitive with OpenAI's models, despite the US's GPU export ban," and attribute it to "innovative use of Mixture of Experts (MoE), Reinforcement Learning (RL) and clever engineering." That is also our reading.

What it changed for us

The honest answer is that it changed what we can study, not what we can build. We were never going to train a frontier model, and we still cannot. But three things that were closed a month ago are open now.

First, the reasoning traces. R1 shows its chain of thought, which means that for the first time we can look at how a reasoning model reaches an answer rather than inferring it from the answer alone. For a group whose work is largely about evaluating and interpreting model behaviour, that is the difference between studying a black box and studying a system.

Second, the distilled models. The report's claim that reasoning transfers to small models by distillation is one we can check directly, because the small models run on hardware we own. We have started running them and will publish what we find once the numbers are stable. That is the kind of experiment that was simply unavailable when the only reasoning model was behind an API.

Third, the price. ORF's report puts o1 at 60 dollars per million output tokens and R1 at 2.19 dollars, roughly a thirty-fold difference. For a non-profit, evaluation budgets are set by inference cost, and a thirty-fold drop turns a study we would have had to sample down into one we can run on the full test set.

What we are setting up

Concretely, three things are on our bench this week. The first is a comparison of the distilled models against their base models on the same tasks, with the chain of thought logged, so we can say how much of the gain comes from the distilled reasoning and how much was already in the base. The report claims the reasoning patterns transfer. We want to know at what size they stop transferring, and whether what transfers is the reasoning or the format of reasoning.

The second is a contamination check. R1 was trained on data we cannot see, like every other model, and its benchmark numbers on mathematics and coding competitions are the ones the report leads with. Before we treat those as evidence of reasoning we want to run problems written after the release date, and we want to publish the problems, so anyone can check whether the next model has seen them.

The third is simply reading the traces. We are collecting chains of thought from R1 on tasks where it gets the wrong answer, and the question is whether the trace shows the model going wrong in a way a person would recognise, or whether the trace and the answer are only loosely connected. If the second, then the visible chain of thought is less of a window than it looks, and that matters for anyone planning to use these traces as a safety signal.

Where the cheap framing misleads

The number that did the damage on 27 January was six million dollars, the reported training cost, set against the more than hundred million dollars OpenAI has said GPT-4 cost. ORF repeats the figure and adds the caveat that it "may not capture full research and development expenses." We want to be more specific about why the comparison is misleading, because the misreading is going to shape funding decisions for a year.

The six million figure, as best anyone outside DeepSeek can tell, is a final training run for the base model, priced at rental rates for the GPUs. It excludes the cluster DeepSeek already owns, the failed runs, the salaries, and the years of prior models the recipe was refined on. The RL stage that turns the base model into R1 is, from the report's description, a much smaller job than pretraining. A base model is a sunk cost that few organisations have paid, and the headline treats it as if it were free. What DeepSeek showed is that the marginal cost of reasoning on top of an existing strong base is low. It did not show that the base is cheap.

There is a second confusion. The market read the release as evidence that less compute is needed. The paper reads to us as evidence that compute is being used more efficiently, which is a different claim with the opposite implication for demand. If reasoning can be added to a base model for a small fraction of what the base model cost, everyone with a base model will buy it, and the inference bill for all those chains of thought will be paid on the same hardware the market just sold.

And a third. DeepSeek built this under export controls on the chips it could legally buy. That is a real engineering achievement and it is also being read as proof that controls do not work. The paper is more consistent with a narrower claim, that controls raise the cost of reaching a given capability without preventing it, which is roughly what controls are for.

What we think comes next

ORF notes that within weeks of the release, a Berkeley group reproduced a small reasoning model for a few hundred dollars, and then for around fifty. The reproduction race has begun and it will run mostly in universities and small labs, because that is who benefits from an open recipe. Our contribution will be on the evaluation side. Someone needs to establish how much of the reasoning is real, how much is pattern-matching on the RL reward, and where the distilled models fall over.

We would also want someone to pin down the cost accounting properly, with the base model, the RL stage, and the distillation stage priced separately. Until that happens the six million number will keep doing work it cannot support, and the next small lab that promises a frontier model on a foundation grant will be quoting it.

Sources

  1. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (arXiv)
  2. Mercer, Spillard and Martin, "Brief analysis of DeepSeek R1 and its implications for Generative AI" (arXiv)
  3. ORF, "Global AI's DeepSeek Moment: Impact and Implications"