OLMo 2 32B and the moment fully open caught GPT-4o mini
Ai2 released a 32B model with public data, code and weights that it says outperforms GPT-3.5 Turbo and GPT-4o mini. What fully open buys a researcher that open weights do not, and how far behind the frontier that still is.
The claim
On March 13 Ai2 released OLMo 2 32B and described it as the first fully open model to outperform GPT-3.5 Turbo and GPT-4o mini. Fully open here has a specific meaning that the rest of the field has slowly drained out of the word open. It means the pretraining data is published, the mid-training data is published, the post-training data and preference data are published, the training code is published, and the weights at every stage, base, SFT, DPO and final instruct, are published. Nothing needed to reproduce the model is withheld.
Ai2 says the model matches or exceeds GPT-3.5 Turbo, GPT-4o mini, Qwen 2.5 32B and Mistral 24B on its academic benchmark suite. We have not run the comparison ourselves, and the benchmarks are the lab's own selection, so take the ordering as Ai2's claim rather than as settled. The part that does not depend on benchmark choice is the recipe, and the recipe is the story.
The recipe
Training ran in three stages. Pretraining used OLMo-Mix-1124, a 3.9 trillion token set drawn from DCLM, Dolma, StarCoder and Proof Pile II, with the 32B model seeing about 6 trillion tokens, which is roughly 1.5 epochs. Mid-training used Dolmino-Mix-1124, 843 billion tokens of higher-quality curated data, with the final checkpoint produced by souping several runs. Post-training followed the Tulu 3 pattern: supervised fine-tuning, DPO on preference data, then reinforcement learning with verifiable rewards using GRPO.
The compute is modest by frontier standards. The run used 160 nodes of Google Cloud H100s, 1,280 GPUs, at about 1,800 tokens per second per GPU, which Ai2 reports as roughly 38 percent model FLOPs utilisation. The blog claims the model was trained at about one third of the cost of Qwen 2.5 32B. That comparison rests on an estimate of Qwen's compute that Ai2 has to infer, so we would read it as a rough ratio.
The training code is a rewritten trainer called OLMo-core, with asynchronous checkpointing, minimal GPU to CPU synchronisation, and support for four-dimensional and higher parallelism. That is infrastructure a lab usually keeps to itself. Releasing it means the 38 percent MFU figure is something another group can try to hit with the same code.
What fully open buys
An open-weights model lets you run it, fine-tune it and probe it. It does not let you answer the questions a researcher most often wants to ask. Was this benchmark in the training data? What did the model look like halfway through training? Which stage of post-training introduced this behaviour? Would a different data mix have changed this? With weights alone, every one of those is a guess.
With OLMo 2 32B all four are experiments. Contamination can be checked by grepping the corpus. The intermediate checkpoints let you watch a capability appear. The SFT, DPO and final weights let you diff behaviour across stages, and the post-training data tells you what each stage saw. A different data mix can be tried on the same code and compared. For interpretability, safety and data-curation research, that is the difference between studying a model and studying a rumour about one.
There is a second-order benefit that matters for a lab like ours. When a result is produced on OLMo, it can be replicated by anyone with the compute, and the replication can vary one factor at a time. A result on a closed model can be replicated only by the lab that owns it, and a result on an open-weights model can be replicated only up to the point where the training data would have to be inspected.
How far behind the frontier
The comparison models tell you the gap. GPT-3.5 Turbo is from late 2022. GPT-4o mini is a small model released in mid-2024, priced for volume rather than capability. Beating them in March 2025 puts fully open somewhere around a year to two years behind the closed frontier, and further behind on the reasoning-model axis, where the comparison in the blog is silent.
That gap is not a reason to dismiss the release. For most research questions about how language models learn, a model that would have been frontier eighteen months ago is more than enough, and the fully open model is the only one on which those questions can be answered rigorously. The gap matters for a different reason: a lot of the field's most interesting behaviours, in reasoning and agentic use, are showing up first in models the open recipe has not caught up to, and studying them on OLMo means waiting.
What we would like to see next is a reasoning-tuned variant with the RL data and the reward code published, so that the RLVR stage can be studied with the same transparency as pretraining. The Tulu 3 recipe was a start. A 32B model with the whole chain public, including the RL, is the thing that would let outside researchers study how reasoning training works rather than inferring it from system cards.
Sources
From the foundation