The estimate

Epoch AI has published a revised version of its 2022 data exhaustion analysis, and the headline is that the total effective stock of public human-generated text is on the order of 300 trillion tokens, with a 90% confidence interval from 100T to 1000T. Under current trends in dataset growth, their 80% interval says that stock will be fully used at some point between 2026 and 2032.

The authors are Pablo Villalobos, Anson Ho, Jaime Sevilla, Tamay Besiroglu, Lennart Heim and Marius Hobbhahn. The 2022 paper had the same basic shape and a more pessimistic tone. What changed in the revision is mostly the treatment of repeated data and of filtered web data, and those two changes are worth examining because they carry most of the conclusion.

Three assumptions that move the date

The first assumption is that models can train for multiple epochs on the same data without much degradation. This was not the consensus in 2022, when one pass over the data was the norm, and it is the main reason the 2024 estimate is less alarming than the original. If repetition works for several epochs, the effective stock is several times the raw stock. If it turns out to work for only two, the exhaustion date moves earlier.

The second is that carefully filtered web data outperforms curated corpora. The 2022 paper drew a hard line between high-quality and low-quality text and worried mostly about the former. The revision drops that distinction and treats filtered Common Crawl as the main resource. Whether you find that reassuring depends on how much you trust the filters, and the paper does not have a way to measure filter quality directly.

The third is the overtraining policy. The authors model developers as profit maximizers who will train small models on far more tokens than compute-optimal scaling would suggest, because inference is where the money goes. Under compute-optimal training the data lasts through 2028. At 5 times overtraining it is fully used by 2027. At 100 times overtraining it is used by 2025, which is next year. So the exhaustion date is partly a choice, and the labs releasing small overtrained models are the ones spending the stock fastest.

Why the number is not the point

The 300T figure is uncertain by an order of magnitude in each direction, and the authors say so. The more useful output is the structure of the argument. Dataset sizes for frontier training runs have grown faster than the internet has, and any two curves with different growth rates eventually cross. The date they cross is uncertain. The fact that they cross is not.

The two growth models the authors use are worth keeping separate. One extrapolates historical dataset growth forward. The other derives dataset size from projected compute, on the assumption that developers will keep training at roughly the token-to-parameter ratios they use today. The two models agree on the rough date and disagree on the mechanism, and the compute-based one is the more fragile, because a shift in how labs allocate compute between pre-training and other stages would move it by years. The last eighteen months of reasoning models trained mostly with reinforcement learning suggest that shift is already happening.

What the paper cannot tell you is what happens when they cross. Its scope is human-generated public text. It does not model private data, licensed corpora, transcribed audio or video, or anything a model generated. Those are exactly the sources labs have been moving toward, so the projection describes the end of one regime rather than the end of progress.

How labs are already responding

The paper itself names three paths beyond 2030: synthetic data, transfer from data-rich modalities, and data efficiency improvements. From where we sit each of those is already in production somewhere. Multi-epoch training is now routine. Synthetic data was the whole premise of the phi models last year. Multimodal pretraining on images and video is standard at the frontier, and the argument that it transfers to text is being tested at scale right now.

Synthetic data is the one we are least sure about. The evidence that it works comes mostly from code and math, where there is a verifier. For open-ended text the risk is that a model trained on outputs of another model inherits that model's blind spots, and we have not seen a convincing measurement of how much of the 300T stock a synthetic corpus can replace before quality drops.

The response we find most interesting is reinforcement learning as a data source. If a model can generate its own training signal by trying things and checking the result, then the relevant constraint is verifiers and compute rather than text. That does not sidestep the wall so much as build a different road, and whether it reaches the same places is an open question.

What we would want measured

The projection would be much more useful with a direct estimate of how many epochs are actually free. A study that trains matched models on one, two, four and eight passes over a fixed corpus, at a scale where the loss curves are trustworthy, would collapse the uncertainty more than anything else in the paper. Some of that exists at small scale. It needs to exist at a size where the answer matters.

The other measurement is the quality of filtered web data against curated text at equal token counts. If filtered Common Crawl really matches books and papers, the 2022 worry about high-quality data goes away. If it does not, the 2024 revision is too optimistic. Either way, the answer is an experiment, and a foundation like ours can run a version of it without a frontier budget.

Sources

  1. Will we run out of data? Limits of LLM scaling based on human-generated data (Epoch AI, 2024)
  2. Will we run out of data? (arXiv 2211.04325)