The DeepSeek R1 shock as an open-science event
An MIT-licensed reasoning model with a detailed paper knocked $589 billion off Nvidia in one session. The part of the story that matters more to researchers is what was published, what was not, and how fast people started replicating it.
The part everyone saw
On January 27 Nvidia's stock fell about 17% and the company lost $589 billion of market value, the largest single-day loss for any company on record. The Nasdaq fell 3.1%, and Arm, Broadcom and Oracle each dropped more than 10%. The trigger was DeepSeek-R1, a reasoning model released a week earlier by a Chinese lab that reported performance comparable to OpenAI's o1 at a fraction of the price.
The training cost figure that drove the panic needs qualification. DeepSeek's V3 technical report gives 2.788 million H800 GPU-hours on a 2,048-GPU cluster, which at a $2 rental rate comes to $5.576 million. That is the cost of the final training run for the base model. Charles Mok's analysis for Stanford's Cyber Policy Center points out it excludes the hardware itself, which DeepSeek reportedly owns, and all prior research and experimentation. Estimates that include those run past a billion dollars. Both numbers are true of different things, and most of the market coverage conflated them.
What was actually published
R1 and R1-Zero are 671B parameter mixture-of-experts models with 37B active, released under the MIT licence. The licence explicitly permits commercial use, modification and derivative works, including distillation to train other models. Alongside them DeepSeek released six distilled dense models, at 1.5B, 7B, 14B and 32B on Qwen bases and 8B and 70B on Llama bases, so that the reasoning behaviour can be studied on a single GPU.
The paper is the more important artifact. R1-Zero applies reinforcement learning directly to the base model with no supervised fine-tuning step, and the authors report that reasoning behaviours emerged from that process alone. R1 adds a small cold-start SFT stage before RL, then a second RL stage and a second SFT stage. The reported numbers are 79.8% on AIME 2024 against o1's 79.2%, 97.3% on MATH-500 against 96.4%, and a Codeforces rating of 2029 against 2061. These are DeepSeek's own evaluations and we have not seen an independent run yet.
What was not published is just as specific. The Hugging Face Open-R1 team listed three gaps on January 28: the reasoning datasets used for SFT and RL were not released, no training code was released, and there is no information about the compute and data scaling trade-offs. So the recipe is described in prose, the weights are free, and the ingredients are missing.
Why this is an open-science story
The reason R1 mattered to us as a researcher rather than as an investor is that a frontier-class reasoning model came with a paper that says how it was trained and a licence that lets anyone try. OpenAI's o1 has neither. For four months the field had been guessing at what makes a reasoning model work. R1's paper offers a testable claim, that pure RL on a base model with verifiable rewards is enough to produce long chains of thought, and the distilled models let a lab with modest compute check pieces of it.
The replication effort started within a day. Open-R1's plan has three steps: reproduce the distilled models by generating reasoning data from R1 itself, reproduce the R1-Zero RL pipeline on curated math, reasoning and code datasets, and then show the full base-to-SFT-to-RL progression end to end. That is exactly the structure of a replication study, and it is being done in the open on GitHub. Whether it succeeds will tell us more about the paper's claims than any benchmark table.
The Stanford piece also reports the allegations. OpenAI has said it has some evidence of distillation from its models, Microsoft researchers observed unusual data extraction through OpenAI's API in the autumn, and the model's censorship of politically sensitive topics is well documented. None of that is settled, and none of it changes what is on Hugging Face. If parts of R1 were built on outputs from a closed model, that is a question about DeepSeek's conduct. It does not make the RL result less checkable.
What we are doing with it
We have started running the 7B and 14B distilled models locally and are looking at whether the chain-of-thought they produce is load-bearing, meaning whether truncating or perturbing it changes the answer. That is a question you cannot ask of a model behind an API that hides its reasoning. It is a question you can ask of an MIT-licensed checkpoint on a single GPU by Thursday.
We are also following Open-R1 rather than trying to duplicate it. A replication is more valuable when it is done once, carefully, in public, than when ten groups each do a quarter of it. What we can add is small-scale ablations on the parts they have not reached yet, especially the claim that no SFT is needed for reasoning to emerge, which is the most interesting and least verified sentence in the paper.
What we think comes next
The market story will fade once people separate the training-run cost from the total cost. The open-science story will not, because the licence does not expire. Within a few months there will either be a public reproduction of R1-Zero's RL result on an open base with open data, or there will be a public failure, and either outcome moves the field more than the original release did.
The thing we would most want someone to try is the smallest possible version. Take a 1.5B base model, a few thousand verifiable math problems, and the GRPO setup from the paper, and see whether anything resembling reasoning emerges. If it does, the claim is real and cheap to check. If it does not, we learn where the threshold is, and that is worth knowing too.
Sources
From the foundation