Tülu 3 opens the post-training black box
AI2 released a complete post-training recipe, data, code and evaluation suite alongside the models. What it documents about supervised fine-tuning, preference tuning and verifiable rewards, and what the decontamination pass revealed about existing open datasets.
What was released
Nathan Lambert, Jacob Morrison, Valentina Pyatkin, Shengyi Huang, Hamish Ivison, Faeze Brahman and sixteen co-authors at the Allen Institute for AI posted the Tülu 3 report on November 22. The models are post-trained from Llama 3.1 base weights at 8B and 70B. The report says they surpass the Llama 3.1 Instruct versions at the same sizes, and that the 70B model passes GPT-4o-mini and Claude 3.5 Haiku on the authors' evaluation suite.
The benchmark claims are the least interesting part. What matters is the list of things that ship with the weights: the full training data mixtures, the data curation tooling, the training code in open-instruct, the evaluation toolkit, and a report long enough to reproduce the run. Post-training has been the least documented stage of building a model. Open base models exist, and the papers describing how they were pretrained are reasonably detailed, but the step that turns a base model into something people talk to has mostly been described in a paragraph.
The three stages
The recipe runs supervised fine-tuning, then direct preference optimisation, then a reinforcement learning stage the authors call RLVR. The SFT mixture is 939,344 prompts. The DPO stage uses 425,145 preference pairs, with variations by model size. Both numbers are large compared to what most open recipes have used, and the report spends much of its length on how the mixtures were assembled skill by skill, with separate targeted subsets for maths, code, instruction following, safety and general chat.
A good example of the data work is the synthetic portion. Rather than generating from a single prompt template, which tends to collapse onto a narrow style, the team conditioned generation on a large set of personas and produced roughly 220,000 maths reasoning instances and 35,000 coding instances that way. The persona is a diversity device. It is a cheap trick, and it is the kind of cheap trick that never appears in a closed lab's model card.
RLVR is the piece with a new name. Instead of a learned reward model, the policy is rewarded when its output can be checked mechanically: a maths answer matches the reference, a constraint in an instruction-following prompt is satisfied. The datasets are GSM8K, MATH and IFEval. The idea is not exotic, and its virtue is that a verifiable reward cannot be gamed the way a learned reward model can. The report treats it as a final polish on the DPO model rather than a replacement for the earlier stages.
What decontamination turned up
The part of the report we keep returning to is the contamination audit. The team checked every candidate training set against their evaluation suite using 8-gram overlap and flagged any dataset where more than 2 percent of instances overlapped with an eval. Several widely used open datasets failed. NuminaMath-TIR had 11.3 percent of instances removed. WildChat's GPT-4 subset lost 5.4 percent. Evol CodeAlpaca lost 3.5 percent. UltraFeedback, which has been the default preference dataset for open DPO work for a year, was contaminated enough that AI2 released a decontaminated version.
Read that list against the leaderboards of the past twelve months. Many open post-trained models used some combination of exactly these datasets, and their scores on the evals in question were reported as measures of alignment quality. Some fraction of that quality was the eval leaking into training. Nobody needed to cheat for this to happen. The datasets were assembled from web scrapes and model outputs that already contained the test items, and no one had run the check.
This is what we mean by post-training being a black box. Malice had little to do with it. Nobody had published a recipe in enough detail for the contamination to be visible, so the field had no way of knowing how much of the reported alignment was undisclosed data.
Development evals and unseen evals
The second design choice we want to flag is the split between development and unseen evaluations. The team picked a set of benchmarks to tune the recipe against, and a separate set it did not look at until the end. Reporting both lets a reader see how much of the improvement is specific to the benchmarks the recipe was optimised for.
This is an ordinary practice in most of machine learning and a rare one in language model post-training, where the eval suite is usually whatever makes the model look best. Combined with the decontamination pass, it means Tülu 3's numbers come with a chain of custody that the numbers it is compared against mostly lack. Whether the 70B model really beats GPT-4o-mini is something we hold loosely, since the comparison was run by the people who built one of the two models. That the comparison was run on decontaminated data with a held-out set is something we can check.
What to do with it
The first thing we would like someone to do is rerun a few popular open post-training recipes on the decontaminated UltraFeedback and report how far the scores move. That is a direct measure of how much of last year's open alignment progress was leakage.
The second is to treat the RLVR stage as a hypothesis rather than a result. Three verifiable domains is a small set, and it is not obvious how far the approach extends past tasks with a checker. The report gives enough detail to test that, which is the whole point of releasing it this way.
Sources
From the foundation