Llama 3 at 15 trillion tokens: how far past Chinchilla can you go?
Meta trained an 8B model on roughly 75 times the compute-optimal token count and says it was still improving. Notes on why that is a rational choice once you count inference, and what the announcement does and does not tell us.
The number that matters in the announcement
Meta released Llama 3 on April 18 in two sizes, 8B and 70B, and buried the most interesting sentence in the training section. The 8B model was pretrained on 15 trillion tokens. Meta's own post states that the Chinchilla-optimal amount of training data for an 8B model is about 200 billion tokens, so the 8B run went roughly 75 times past that point. The post says both the 8B and 70B models continued to improve log-linearly all the way to 15 trillion tokens.
That is the claim we want to sit with. Chinchilla, the 2022 DeepMind result, told the field that for a fixed training compute budget you should scale parameters and tokens together, and that most models at the time were far too large for the data they saw. Meta has taken the opposite tack and kept a model small while feeding it an enormous corpus. The 15 trillion token dataset is seven times the size of the one used for Llama 2, with four times as much code and over 5 percent non-English text across more than 30 languages.
Compute-optimal is a statement about training only
Chinchilla answers a narrow question. Given this many FLOPs for training, which combination of model size and token count gives the lowest loss? It says nothing about what happens after training ends. A model that is served to millions of users for a year spends far more compute on inference than on its own training run, and the cost of every forward pass scales with parameter count, not with how many tokens the model saw.
So the accounting changes once inference is in the budget. A smaller model that has been trained well past its compute-optimal point is more expensive to produce than the compute-optimal model of equal quality, but it is cheaper to run for every token it ever generates. Meta says this directly in the post, noting that smaller models are preferable for inference efficiency even when larger models could match their performance with less training compute. The training overspend is a one-time payment to lower a recurring bill.
This is why we think of the two regimes as compute-optimal versus cost-optimal. Compute-optimal minimises training FLOPs for a target loss. Cost-optimal minimises total spend over the life of the model, training and serving together. The second objective pushes hard towards small models trained long, and the further inference volume grows, the harder it pushes. Meta is one of the few organisations that serves models at a scale where the difference is worth tens of thousands of GPUs.
What log-linear improvement does and does not mean
The phrase log-linear improvement deserves care. It means that each doubling of tokens bought roughly the same absolute gain in whatever metric Meta was tracking. It does not mean the gains were large. Going from 7.5 trillion to 15 trillion tokens costs as much compute as the entire first 7.5 trillion, and buys one more increment of the same size as the previous doubling. At some point the increment stops paying for itself, and the post does not tell us where Meta thinks that point is for an 8B model, only that it had not been reached at 15 trillion.
The other missing piece is what was measured. The post reports downstream benchmark results and human evaluations across 1,800 prompts in 12 categories, compared against Claude Sonnet, Mistral Medium and GPT-3.5. It does not give a loss curve against token count. Without one, log-linear is a description we have to take on trust, and we would want to see the curve before treating 15 trillion as a calibrated recommendation rather than a data point.
The engineering that made it affordable
A 15 trillion token run at 8B parameters is not exotic compute by frontier standards, but Meta's infrastructure numbers explain how the larger runs were possible at all. The post describes two custom-built clusters of 24,000 GPUs, training utilisation above 400 TFLOPS per GPU, and effective training time above 95 percent thanks to automated error detection and maintenance. Meta says the overall training efficiency improved about three times over Llama 2.
The tokenizer change also matters more than it looks. Llama 3 moved to a 128K token vocabulary, which Meta says encodes text with up to 15 percent fewer tokens than the Llama 2 tokenizer. Fewer tokens per document means each of the 15 trillion tokens carries more text, and at inference time it means fewer forward passes per response. A vocabulary change is a cheap way to buy a small efficiency win on both sides of the ledger, and it is telling that a lab optimising for serving cost took it.
What we would want to see next
The obvious experiment is the one Meta has the data for and has not published. Take the checkpoints of the 8B model at 1, 2, 4, 8 and 15 trillion tokens, evaluate each on the same held-out set, and put the curve next to the marginal training cost at each point. That single figure would let anyone with a serving forecast work out where their own cost-optimal point lies, rather than inferring it from a sentence in a launch post.
The post also mentions a model above 400 billion parameters still in training on the same data. If Meta publishes per-token curves for both sizes, we would finally have a public estimate of how the overtraining benefit changes with scale, which is the number the field has been guessing at since Chinchilla. We would trade most of the benchmark tables in the announcement for that one plot.
Sources
From the foundation