How many epochs is too many? Scaling data-constrained language models
Muennighoff and colleagues trained 400 models to find out what repeating data costs. Up to about four epochs it costs almost nothing, and the fitted law says why. Reading notes on a paper that quietly resets the assumptions behind the running out of data worry.
The question Chinchilla did not answer
The Chinchilla scaling result told everyone to grow data and parameters together, and it assumed you always have more unique text to grow into. That assumption is starting to bind. High quality public text is finite and the compute budgets of the largest runs are approaching the point where a compute optimal token count exceeds what anyone has. The standard response has been to worry. The paper Niklas Muennighoff, Alexander Rush, Boaz Barak, Teven Le Scao, Aleksandra Piktus, Nouamane Tazi, Sampo Pyysalo, Thomas Wolf and Colin Raffel posted last week measures instead.
They trained around 400 models from 10 million to 9 billion parameters, up to 900 billion tokens, on C4 and OSCAR with the GPT-2 architecture and tokenizer. The variable is how many times the same unique tokens are seen. Hold compute fixed, shrink the unique data, and repeat it to fill the budget. Then see how loss responds.
Four epochs are close to free
The headline finding is that up to four epochs of repeated data produce negligible change in loss compared to training on the same number of unique tokens. In the comparison they report, a model trained for four epochs came in about 0.5 percent higher on validation loss than a single epoch model at the same compute. That is small enough to be within the range where you would take it in exchange for not having to find three times as much data.
The decay after that is gradual rather than a cliff. Returns diminish steadily, and by around 16 epochs they are very limited. At 40 epochs the fitted law predicts almost nothing from further repetition, and adding compute on top of heavily repeated data also decays to zero value. The number to remember is four, with a soft shoulder out to somewhere in the teens.
The fitted law
The paper's contribution beyond the measurement is a scaling law that takes repetition into account. The trick is to replace raw data with effective data. If U_D is the number of unique tokens and R_D is the number of times they are repeated beyond the first pass, the effective data is D' = U_D + U_D times R_D* times (1 minus e to the power of minus R_D over R_D*). The constant R_D* acts like a half life. Repeated tokens lose roughly 63 percent of their value by the time repetition reaches R_D*.
The same form applies to parameters, with an effective parameter count N' built from the unique data and a constant R_N* that discounts excess parameters when there is not enough data to use them. Fitting on 182 runs gave R_D* about 15.4 and R_N* about 5.3. Plugging both into a Chinchilla style loss gives L = 521 over N' to the 0.35, plus 1488 over D' to the 0.35, plus 1.87. The shape of the curve, gentle for the first few epochs and then flattening, falls out of the exponential in D'.
One consequence is a different allocation rule. In the data constrained regime the paper's efficient frontier says to spend extra compute on more epochs rather than on more parameters. Chinchilla's equal scaling advice was for unlimited data and it does not carry over.
Code and filters
Since the whole worry is about running out of natural language, the authors also tested two ways of stretching the supply. Mixing in Python code up to half of the training data showed no degradation on natural language tasks, which they read as doubling the effective token pool for free. That is a surprising amount of headroom and it is the result we would most want replicated with a different code corpus.
The filtering experiments were more mixed. Perplexity filtering helped on the noisier OSCAR data, though the filtered data then has to be repeated more to fill the budget. Deduplication of the already clean C4 did not improve downstream performance in their setup, which runs against the intuition that duplicates are always bad. The intuition might still be right for web scale crawls with heavy near duplication, but it was not visible on C4.
What this does to the data wall
The running out of data estimate that gets quoted assumes each token is used once. This paper multiplies the usable budget by roughly four before any loss shows, and by more if you are willing to accept a small hit. Add the code result and the ceiling moves by another factor of two. The supply is still finite, and the date at which the constraint bites by enough to change what a lab should plan for this year.
The gap we see is scale. Nine billion parameters and 900 billion tokens is a serious sweep for an academic group and small next to the largest runs. Whether R_D* stays near 15 for a model ten times larger is exactly the thing the law is being asked to predict, and exactly the thing the data cannot yet confirm. The models and datasets are public, so someone with a larger budget could extend the frontier and find out.
Sources
From the foundation