The Common Pile: can you train a competitive model on licensed text only?
EleutherAI and collaborators released 8TB of public domain and openly licensed text and trained two 7B models on it. The result lands near Llama 2 7B on most benchmarks, and the places where it falls short tell you exactly what the open web contributes.
What was released
The Common Pile v0.1 is an 8TB corpus assembled from 30 sources that are either in the public domain or carry an open licence, and the paper that describes it reports two 7 billion parameter models, Comma v0.1-1T and Comma v0.1-2T, trained on one and two trillion tokens of it. The authors say both models reach performance competitive with models trained on unlicensed text at similar compute, and they name Llama 1 and Llama 2 7B as the comparison points. The dataset, the collection code, the training mixture and the checkpoints are all public.
The collaborator list is long. EleutherAI led, with contributors from the University of Toronto, Vector Institute, Hugging Face, the Allen Institute for AI, Teraflop AI, Cornell, MIT, CMU, Lila Sciences, poolside, the University of Maryland and Lawrence Livermore National Laboratory. Twenty seven people are on the paper.
What is in the mix
The sources sort into a handful of domains. Scientific text comes from peS2o, PubMed Central and arXiv. Government and legal text comes from the US Government Publishing Office, Google Patents public data, the Caselaw Access Project and Court Listener, UK Hansard and Regulations.gov. Books come from the Biodiversity Heritage Library, pre-1929 scans via Internet Archive and HathiTrust, the Library of Congress digitised books and Project Gutenberg. There are open textbooks from the Directory of Open Access Books, PressBooks, OERCommons and LibreTexts, discussion text from StackExchange, GitHub Archive and Ubuntu IRC, code from the openly licensed part of Stack v2, plus Wikimedia wikis, a Creative Commons subset of Common Crawl, and transcripts from over a million Creative Commons YouTube videos.
The licensing standard is stricter than most people expect. The authors only took data where they were confident the licence was set by the actual copyright holder, to avoid what they call licence laundering. They excluded OpenAlex and YouTube Commons because the metadata was unreliable. They refused collection level licences like ODC-By that do not flow down to individual documents. They rejected synthetic datasets generated by models that were themselves trained on unlicensed text. For the web crawl, they manually checked the 1,000 domains contributing the most content. Code from Stack v2 was included only where licence identification came through the Software Heritage Foundation and ScanCode.
The processing after that is fairly standard. English language filtering with FastText, a quality classifier adapted from DataComp-LM for the web portion, an OCR error detector based on unigram likelihood for the scanned books, toxicity filters trained on Jigsaw, PII redaction, and fuzzy deduplication at 90 percent 20-gram overlap. The mixing weights were chosen by training a small model per source at 28 billion tokens and up-weighting the sources that did well, with a cap of six repetitions of any source in the 1T run.
Where Comma lands, and where it does not
On the 1T budget, the paper compares Comma to Llama 1 7B. Comma does well on ARC-C, MMLU, HumanEval and MBPP. It lags on HellaSwag and PIQA. On the 2T budget, the comparison set is Llama 2 7B and OLMo-7B-Twin-2T, and the pattern is the same. Comma is strong on MMLU, SIQA, ARC-E and coding, and weaker on the commonsense suite of HellaSwag, PIQA and CommonSenseQA. The exact percentages are in the appendix and we would encourage anyone who wants to cite the model to read them rather than the summary, because the gap is benchmark specific rather than uniform.
The authors have a concrete explanation for the commonsense gap, and it is the most useful finding in the paper for anyone building corpora. Their analysis suggests HellaSwag performance depends most on coverage of personal blogs, tutorials, hobbies and sports. Those are exactly the domains you cannot get under a verifiable open licence at scale. Government documents, court opinions, patents and scientific papers are available in bulk. Someone writing about their weekend hike is not. So a licensed corpus is skewed toward formal registers, and the benchmarks that reward informal, everyday knowledge notice.
The controlled comparison makes the same point from the other side. With 1.7B models trained on 28B tokens, Common Pile beat the other open licence corpora, KL3M, OLC and Common Corpus, on every benchmark, and matched or beat the original Pile and OSCAR on most. FineWeb, which is unlicensed web text, still did best on most tasks, though Common Pile won on MMLU and ARC. The authors attribute FineWeb's edge to its much larger pool of raw text, which lets you filter aggressively and keep only the top slice. A licensed corpus has no such slack.
What the 2T run tells you about the ceiling
The 2T model reused the 1T mixture, which meant roughly 16 passes over some sources. The authors call this out as probably suboptimal. That is the honest version of the ceiling on this approach right now. Eight terabytes sounds large, but after filtering, deduplication and mixing, the amount of high quality diverse text is small enough that a 2T token run is already recycling. Qwen3 8B, trained on 36T tokens, is included as an upper reference and sits well above every budget matched model. The licensed corpus is not close to that regime and the paper does not pretend otherwise.
There are also caveats the authors state that we want to repeat. Licence laundering is hard to catch exhaustively. The metadata is frozen as of late 2024. Public domain documents can quote in-copyright material. None of that undermines the release, but it does mean a downstream user should treat the licence guarantee as a best effort with a documented method rather than a legal certainty.
What we would try next
The obvious experiment is to fill the informal register gap without breaking the licensing rule. The CC YouTube transcripts are the closest thing in the corpus to everyday spoken language, and we would like to see an ablation that up-weights them and measures HellaSwag and PIQA specifically. If the gap closes, the commonsense story is about register and not about volume. If it does not, the missing ingredient is something else.
The second thing we would want is a re-run of the 2T model with a mixture rebalanced for two trillion tokens rather than one, since the authors already suspect the repetition hurt. Both experiments are cheap relative to the original run, both use only the released artefacts, and both would tell us whether the remaining distance to unlicensed models is a data availability problem or a data curation problem. Those have very different implications for whether this approach scales.
Sources
From the foundation