What model flow means

Ai2 released OLMo 3 on November 20 and the phrase they chose to describe it is the model flow. The idea is that a language model consists of every stage that produced the weights file, the datasets at each stage, the checkpoints between stages, and the code that ran them. Ai2 is publishing all of it. Most releases that call themselves open give you the last item on that list and nothing before it.

The family is two sizes. At 7B there are Base, Instruct, Think and RL Zero variants. At 32B there are Base and Think, with 3.1 updates adding a 32B Instruct. Ai2 positions 32B as the size a research group can fine-tune on hardware it owns while still being competitive, and the training ran on a 1,024 H100 cluster.

The stages, each with its own data

Pretraining has three stages and each has a named dataset. The full Dolma 3 pool is about 9.3 trillion tokens of web, scientific PDFs processed with Ai2's olmOCR, code, math and encyclopedic text. Dolma 3 Mix, the actual pretraining set, is 5.9 trillion tokens with the code and math share raised and extensive deduplication and decontamination. Dolmino, the mid training set, is 100 billion tokens sampled from a 2.2 trillion pool of math, science, code, instruction and reading comprehension data, and it includes reasoning traces. Longmino, about 50 billion tokens from a 639 billion token pool of long documents, extends context, and Ai2 reports the base model holds quality up to about 65K tokens on RULER.

Post training follows a three stage recipe of supervised fine-tuning, DPO and reinforcement learning with verifiable rewards, on a dataset suite called Dolci with a separate mix per stage and per target, so Think and Instruct get different data. Every one of these datasets is downloadable, and Ai2 says without license restrictions. Simon Willison's write-up notes the pool is still crawled web text rather than exclusively licensed content, with robots.txt and paywalls respected, which is the honest framing.

The token efficiency claim

The number Ai2 leads with is that Olmo 3 Think 32B trains on roughly 6x fewer tokens than the models it is compared against, and is within a few points of them. On MATH it reports 96.1 against 96.7 for Qwen 3 VL 32B Thinking. On IFEval it reports 89.0, ahead of every comparison listed. On HumanEvalPlus it is 91.4, within a point of the DeepSeek R1 distilled 32B. On BigBenchHard 89.8. The base 32B model posts 66.5 on HumanEval, 81.0 on DROP and 80.5 on GSM8k.

We take the 6x figure as an efficiency claim about the pipeline rather than a claim about a single trick, and that is what makes it checkable. Every token in the 5.9 trillion mix is published, so anyone can ask whether the efficiency comes from the data selection, the mid training reasoning traces, or the RLVR stage, by training from the intermediate checkpoints rather than from scratch. With a Qwen or DeepSeek weights release you can only compare final models, and if one wins you cannot say why.

Grouped bar chart comparing Olmo 3-Think 32B against Qwen 3 32B, Qwen 3 VL 32B Thinking, Gemma 3 27B, and DeepSeek R1 Distill 32B across four benchmarks. MATH: 96.1, 95.4, 96.7, 87.4, 92.6. IFEval: 89.0, 86.5, 85.5, 85.4, 78.7. HumanEvalPlus: 91.4, 91.2, 90.6, 79.2, 92.3. BigBenchHard: 89.8, 90.6, 91.1, 82.4, 89.7.
Olmo 3-Think 32B against its comparison set, from Ai2's own release table. Trained on roughly 6x fewer tokens, it lands within a few points across all four.

What you can study that weights alone do not allow

Three things stand out. First, the RL Zero pathway. Ai2 publishes checkpoint series for reinforcement learning run directly on the base model in four domains, math, code, instruction following and general chat, precisely so that RL algorithms can be benchmarked against a common starting point with known data. Arguments about whether RL teaches reasoning or elicits it need exactly this, a base model whose pretraining data you can search.

Second, OlmoTrace. In the Ai2 Playground you can take a model response and trace spans of it back to matching documents in the training data in real time. Simon Willison tried it on a biographical query and found the matches not very relevant to what he asked, which we would expect for a feature that matches text rather than meaning. It is still the only place we know of where you can point at a generated sentence and see the training text it resembles.

Third, the intermediate checkpoints. Base, mid trained, long context extended, and every post training stage on both the Think and Instruct paths. If a behaviour appears in the Think model you can bisect the pipeline to find the stage that introduced it. That is the experiment interpretability people keep wanting to run on production models and never can.

The costs of a thinking model you can inspect

Simon Willison's practical note is that the 32B Think model overthinks. His SVG generation request spent over 14 minutes reasoning. That is a real cost of a model whose reasoning traces are visible, and it is also data. Someone can now take that trace, find the point where the model had the answer and kept going, and look at what the mid training reasoning traces taught it about when to stop.

The infrastructure release came with it, Olmo-core for distributed training, Open Instruct for post training, the Rust based datamap-rs and duplodocus tools for cleaning and fuzzy deduplication, decon for test set removal, and the OLMES evaluation suite. Ai2 reports 8x faster SFT and 4x faster RL throughput from in flight weight updates and continuous batching. What we would want next is a third party training a 7B model from Dolma 3 Mix on independent hardware and reporting whether the published recipe reproduces the published base model. Until someone does that, fully open is a promise about reproducibility rather than a demonstration of it.

Sources

  1. Ai2, Olmo 3: Charting a path through the model flow to lead open-source AI
  2. Simon Willison, Olmo 3 is a fully open LLM