Pythia and the case for publishing checkpoints, not just weights
EleutherAI trained 16 models from 70M to 12B parameters on identical data in identical order and released 154 checkpoints for each. A look at the design, and at the questions about memorisation and training dynamics that only this kind of suite can answer.
What was released
Pythia is a suite of 16 autoregressive language models from EleutherAI, ranging from 70M to 12B parameters, all trained on the Pile in exactly the same order. Eight sizes were trained on the standard Pile and the same eight again on a deduplicated version. For each of the 16 models there are 154 saved checkpoints, and the paper ships tooling to reconstruct the exact training dataloader so you can recover which batch the model saw at any step.
The checkpoint schedule is what makes this different from a normal release. Checkpoints are saved every 1,000 iterations, which at a batch of 1,024 sequences of 2,048 tokens is roughly every 2.1 billion tokens, giving 144 evenly spaced saves across the roughly 300 billion token run. On top of that there are log-spaced checkpoints early in training at steps 1, 2, 4, 8, 16, 32, 64, 128, 256 and 512. The paper's comparison table shows how unusual this is. Most public suites offer one checkpoint per model, OPT offered two, and BLOOM eight.
The authors are explicit that they did not chase benchmark performance. They chose a batch size larger than usual for the smaller models, used fully dense layers rather than the sparse and dense alternation in GPT-3, and trained the deduplicated models for about 1.5 epochs since the deduplicated Pile is about 207 billion tokens. They report that the models are competitive with OPT and GPT-Neo at similar sizes, which is the bar they set, and no higher.
Why a suite and not a model
Almost every scaling result we have compares models that differ in more than size. Different data, different data order, different tokenisers, different hyperparameter sweeps. When a behaviour appears at 6B and not at 1B, you cannot tell whether size caused it or whether one of a dozen uncontrolled variables did. Pythia removes those variables. Every model saw the same tokens in the same sequence, so a difference between the 1B and the 6.9B checkpoint at step 50,000 is a difference attributable to width and depth.
The same holds across time. Because the dataloader is reproducible, you can ask what a model at step 40,000 has seen and not seen, and then test whether it knows something that first appears in the data at step 60,000. That question is unanswerable for any model that ships only final weights.
Memorisation as a Poisson process
The first case study tests whether training order affects memorisation. The natural guess is that sequences seen late in training are memorised more, since the model has had less time to overwrite them. The authors measure memorisation of training sequences across the run and find that the count of memorised sequences per batch is fit very well by a Poisson point process, which implies that where a sequence falls in training has little effect on whether it gets memorised. The rate is roughly uniform.
This is a result you can only get with a reproducible dataloader and a dense checkpoint series, because it requires knowing exactly which sequences were in which batch and measuring at many points. It also has a practical consequence the paper draws out. If you wanted to protect a particular document from memorisation by putting it early or late in the run, that would not work.
Term frequency and gender bias
The second case study follows Razeghi and colleagues in asking whether few-shot performance on arithmetic and question answering tracks how often the relevant terms appeared in pretraining. With Pythia the authors can count term frequencies in the data actually seen up to a given checkpoint. They find that the correlation between frequency and accuracy is an emergent property of the larger models. Smaller models do not learn the tasks regardless of how frequent the operands are, and the paper reports a phase change after 65,000 steps, about 45 percent of the way through training, after which models of 2.8B parameters and above start to show the correlation while it stays largely absent in smaller ones.
The third study is an intervention rather than an observation. The authors take the deduplicated 70M, 410M, 1.4B and 6.9B models and retrain the last 7 percent or 21 percent of the run with morphologically masculine pronouns swapped to feminine, so the models see identical data until the swap begins. They then measure WinoBias and the CrowS-Pairs gender subset. Swapping pronouns for the last 21 percent of training reduces the measured bias at all sizes, and the effect grows with model size, which the authors read as larger models learning more complex occupation and pronoun relationships and therefore responding more to the changed statistics. Because the counterfactual model is a branch off the same run, the only difference is the pronoun frequencies.
What this changes for the rest of us
The practical lesson is that the cost of saving checkpoints is tiny compared to the cost of training, and the scientific value of the series is enormous. A single set of final weights supports evaluation. A checkpoint series with a reproducible dataloader supports experiments about learning itself, which is the part of the field with the least evidence and the most speculation.
What we would try first with the suite is a memorisation study on a deliberately inserted canary set, using the dataloader to know exactly when each canary appears, and a comparison of the 1.4B and 12B trajectories on a task where the final scores match to see whether they got there the same way. Both are cheap now. Neither was possible a month ago.
Sources
From the foundation