FineWeb: what it takes to decant the web into 15 trillion good tokens
Hugging Face released a 15 trillion token pretraining corpus with an ablation for every filtering and deduplication choice, and a 1.3 trillion token subset selected by a classifier trained on LLM judgements of educational quality. Reading notes on why data curation is now a research field with its own results.
What was released
FineWeb is 15 trillion tokens of English text drawn from 96 Common Crawl snapshots, released by Guilherme Penedo, Hynek Kydlíček, Loubna Ben Allal and colleagues at Hugging Face, with the paper posted to arXiv on June 25. FineWeb-Edu is a 1.3 trillion token subset. The data are public, the processing code is public, and so are all the models trained for the ablations. That last part is what makes this a paper rather than a download.
The claim is that models trained on FineWeb beat models trained on RefinedWeb, C4 and the other public web corpora on a suite of eight benchmarks, and that FineWeb-Edu beats everything on knowledge and reasoning tasks like MMLU and ARC. The evidence for each processing decision is an ablation, which is a small model trained on the data with and without the step, so the reader can see what each filter bought.
The pipeline, step by step
Processing starts with a URL blocklist. Text is extracted from the raw WARC files with trafilatura rather than taken from Common Crawl's own WET text files, and the ablation shows the WARC route produces better data. A fastText language classifier keeps English at a 0.65 confidence threshold. Then come the quality filters from Gopher's MassiveText, which remove documents with too few words, too many symbols, or too little of the structure real prose has.
The team adopted the C4 filters next and then went looking for their own. They tested more than 50 heuristics and kept three. Documents where 0.12 or fewer of the lines end with punctuation are dropped, which removes about 10 percent of tokens. Documents where the characters in duplicated lines make up 0.1 or more of the text are dropped, removing about 12 percent. Documents where 0.67 or more of the lines are short are dropped, removing about 4 percent. Each of those thresholds was picked by looking at where the distribution of good and bad data separated, then checked with a training run.
The deduplication surprise
The result we did not expect concerns deduplication. The conventional view was that more deduplication is better, so the first attempt deduplicated across all 96 snapshots at once using MinHash. Performance got worse. When the team looked at what global deduplication had kept from the older snapshots, it turned out to be lower quality than what it had removed, because the documents that survive across many crawls are often the boilerplate that never changes, while the good pages had been dropped as duplicates of newer copies.
The fix was to deduplicate each snapshot on its own. The MinHash configuration is 112 hash functions in 14 buckets, targeting a 75 percent similarity threshold. Per-snapshot deduplication leaves some cross-snapshot duplicates in the corpus and the ablations show that is fine, or better than fine. We take this as a warning about the whole field: a step everyone agreed was good was hurting, and nobody knew because nobody had run the experiment at scale with the results published.
FineWeb-Edu and the classifier
The educational subset is built from a different kind of signal. The team had Llama-3-70B-Instruct score around 460,000 web pages on educational quality, then trained a small classifier on those annotations, then ran the classifier over all 15 trillion tokens. Pages scoring 3 or above on the scale were kept, and that is the 1.3 trillion token FineWeb-Edu.
The ablations here are the strongest in the paper. On MMLU and ARC the Edu subset outperforms the full FineWeb and every other public corpus, and the paper says it matches the performance of larger unfiltered datasets with roughly ten times fewer tokens. Every quality heuristic in the main pipeline is a guess about what good text looks like. The classifier replaces the guess with a model's judgement, and the model's judgement turns out to be a better filter than any rule the team wrote by hand.
The obvious worry is circularity. A classifier trained on one language model's taste, used to select data for training the next language model, will reproduce that taste. The gains on MMLU may partly reflect that MMLU rewards textbook-like text and the classifier was asked to find textbook-like text. The paper does not resolve this and we do not think it can be resolved without a downstream evaluation that the classifier was not designed around.
Why this is now a field
Five years ago the pretraining corpus was a paragraph in an appendix. FineWeb makes it a paper with a methods section, ablations, released artifacts and a negative result. The negative result on global deduplication alone would have justified publication, because it overturns a default that every group had adopted without measuring. The filtering thresholds are the kind of number that only exists if someone did the experiment.
What we want next is for other groups to rerun the ablations on their own model sizes and report where the conclusions change. The FineWeb ablations are on small models by necessity, and a filter that helps at 1.8 billion parameters may not help at 70. The second thing is a classifier trained on a different judge, so we can see how much of FineWeb-Edu's gain is educational quality and how much is Llama-3's opinion of it. The data and code are there. It is a matter of compute and someone deciding it is worth publishing.
Sources
From the foundation