RedPajama: reproducing a training set from a paper's recipe
Together and partners rebuilt the 1.2 trillion token LLaMA training mix from the description in the paper. The token counts land close to the original on most slices and 41 percent short on GitHub, which is a measure of how much a recipe leaves out.
The artifact that did not ship
When Meta published LLaMA in February, the paper described a training set assembled entirely from public sources, and released the data to nobody. The data did not exist as an artifact at all. It existed as a few paragraphs of recipe. On April 17 Together, with Ontocord.ai, ETH DS3Lab, Stanford CRFM and Hazy Research, with support from LAION, EleutherAI and the Oak Ridge Leadership Computing Facility, released RedPajama-Data, a 1.2 trillion token reconstruction of that recipe, with the processing scripts on GitHub and the data on Hugging Face.
We want to write about this as a replication rather than as a release, because the interesting content is in the gap between what the paper said and what the team had to decide. A recipe that says CommonCrawl filtered with CCNet and a Wikipedia classifier is enough to know what was done. It is not enough to do it again and get the same corpus. RedPajama is the first public measurement of how wide that gap is.
The seven slices and how they compare
The LLaMA paper listed seven sources with token counts. RedPajama reports its own counts next to them. CommonCrawl came in at 878 billion tokens against 852 billion in the paper. C4 at 175 billion against 190. Books at 26 billion against 25. ArXiv at 28 billion against 33. Wikipedia at 24 billion against 25. StackExchange at 20 billion against 27. And GitHub at 59 billion tokens against 100 billion in the paper.
Five of the seven land within about 10 percent of the target. The two that miss are C4, which is a fixed public dataset and so the difference is presumably tokenisation and cleaning, and GitHub, which comes in 41 percent short. The team wrote that they tuned quality filters to roughly match the reported counts, which is an honest description of what replication from a recipe looks like. You adjust the knobs you were not told the settings of until the output looks like the number in the table.
Where the recipe was thin
Each slice needed a decision the paper did not make for them. For CommonCrawl, the LLaMA description was the CCNet pipeline plus a linear classifier that selects for Wikipedia-like pages. Which crawls, how many, what threshold on the classifier, and what deduplication scope are all choices, and each one moves the token count and the content. RedPajama published its processing scripts, which is more than the original did.
GitHub is the clearest case. The paper describes license filtering and quality heuristics. Which licenses count as permissive, how forks and vendored copies are handled, what line-length and alphanumeric thresholds define quality, and which snapshot of the GitHub corpus you start from will change the yield by a factor of two without any of them being wrong. The 59 versus 100 billion gap is what happens when a reasonable team makes reasonable choices that differ from another reasonable team's. It is also, we suspect, the slice where the resulting model will differ most, because code is where filtering thresholds bite hardest.
ArXiv, Books, Wikipedia and StackExchange were described as boilerplate removal and deduplication, and those came in close, which suggests that when the source is a bounded collection the recipe underdetermines less. The hard cases are the open-ended crawls where the filter defines the dataset.
Data is the artifact
The reason this matters beyond LLaMA is that the paper is treated as the reproducible object in our field and it is not. A model paper gives you an architecture and a training curve. The architecture is a few hundred lines of code that anyone can retype. The curve depends on the data, and the data is described in a paragraph. If two groups following the same paragraph get corpora that differ by 40 percent on one slice, then the paper is not sufficient to reproduce the result, and the thing that would make it sufficient is the dataset itself, or at minimum the exact filter code and source snapshots.
RedPajama makes that argument by doing the work. The scripts are public, so the next person does not have to guess the thresholds again. Whether the resulting corpus produces LLaMA-quality models is the open question, and the team says training is next, with base models trained through the INCITE programme at Oak Ridge and instruction tuning using OpenChatKit data. That model is the actual test of the replication. Matching token counts is necessary and not sufficient, and we would not be surprised if the GitHub slice shows up as a capability gap on code benchmarks.
What we would want from the next paper
If you are writing a model paper and you cannot release the data, release the filter code with its parameters, the list of source snapshots by identifier, and a hash of the final corpus. That costs nothing in competitive terms that the recipe paragraph does not already give away, and it lets a replication check itself rather than tuning to a table. RedPajama had to reverse engineer numbers that could have been a config file.
And for the replication itself, the experiment we want is an ablation on the GitHub slice. Train two otherwise identical small models, one on the 59 billion tokens and one on a version filtered to a different threshold that lands nearer 100 billion, and see whether the code scores move. That would tell us whether the underspecified part of the recipe was the important part.
Sources
From the foundation