Pythia: the model suite built to be studied, not deployed
EleutherAI has released 16 language models from 70M to 12B parameters, each with 154 checkpoints, all trained on the same data in the same order. The benchmark scores are beside the point. For the first time anyone can watch a model learn.
What was released
This week Stella Biderman, Hailey Schoelkopf and eleven coauthors at EleutherAI posted the Pythia paper and the models it describes. There are eight sizes from 70M to 12B parameters, each trained twice, once on the Pile at about 300B tokens and once on a deduplicated version of the Pile at about 207B tokens. That makes 16 models. For each one there are 154 checkpoints, one at initialisation, a set of log spaced checkpoints early in training, and then one every 1,000 steps. Everything is Apache 2.0, the models are on the Hugging Face Hub, and the training code, analysis code and data loaders are on GitHub.
The key design decision is that all 16 models saw exactly the same data in exactly the same order. Batch size is 1,024 sequences of 2,048 tokens at every scale. The architecture is the same at every scale too, including parallel attention and feedforward blocks, which the received wisdom says hurts small models. The authors kept it anyway, because a suite where the small models differ from the large ones in ways other than size cannot answer questions about size.
Why the controlled trajectory matters
If you want to study how a language model develops, you need to see it at multiple points in training. If you want to study how a behaviour changes with scale, you need models at multiple sizes that differ in nothing but scale. Before Pythia, neither existed in public. Open models were released as a final checkpoint, trained on undisclosed or unordered data, with architectural tweaks at each size. Comparing GPT-2 small to GPT-2 XL confounds size with everything else.
The Pythia suite removes the confounds. Any difference between the 410M model at step 50,000 and the 2.8B model at step 50,000 is due to parameters, because the data seen is identical down to the batch. Any difference between step 50,000 and step 100,000 of one model is due to the next 50,000 batches, and you can look at exactly which sequences those were. That is what makes it a scientific instrument rather than a product.
Three case studies
The paper demonstrates the instrument with three studies. The first is memorisation. A natural hypothesis is that sequences seen early or late in training are memorised more, either because early data shapes the weights or because late data is fresh. The authors checked which training sequences the models can reproduce, and found no such structure. Memorised sequences are spread through training at a rate consistent with a Poisson point process. Position in the data order does not predict memorisation.
The second is the effect of term frequency on few shot performance. The finding here is that there is a phase change around step 65,000. Before it, how often a term appeared in pretraining has little relationship to how well the model handles tasks involving that term. After it, and only for models of 2.8B parameters and above, task accuracy starts to track pretraining term frequency. Smaller models never develop the correlation. This is the kind of result you cannot get without both the checkpoints and the scale ladder.
The third is an intervention. The authors took models partway through training and retrained the final 7 percent or 21 percent of the run with gendered pronouns swapped in the data. Stereotypical gender bias on the evaluation dropped. That demonstrates that a modest late training data change can move a measured bias, and it is a template for the kind of controlled intervention the suite makes cheap.
A public good, with the usual caveats
The models are deliberately not competitive on benchmarks. The paper is clear that they were designed for analysis rather than deployment, and the parallel attention decision alone probably costs the small models some accuracy. If you want the best 1B model to ship, look elsewhere. If you want to know what a 1B model looked like at step 10,000, there is now exactly one place to look.
The other caveat is that everything rests on the Pile, so any finding is a finding about models trained on that particular corpus. The deduplicated variant helps, because it lets you separate effects of duplicated data from effects of the data itself, but it is still one distribution.
What we expect to see over the next year is a stream of papers that would not have been possible in March. Studies of when specific capabilities appear during training. Studies of which training examples are responsible for which outputs. Replications of scaling claims with the confounds removed. The experiment we would run first is to repeat the memorisation study on the deduplicated suite and check whether the Poisson result survives when no sequence is seen twice. If it does, memorisation is about the sequence itself rather than its repetition, and that would change how we think about filtering training data.
Sources
From the foundation