Why this report is different

Meta released the Llama 3.1 models on July 23 and the technical report on July 31. The largest model is a dense 405 billion parameter transformer trained on 15.6 trillion tokens with a 128K context window, and the blog post says it is the first Llama trained on more than 16,000 H100 GPUs. Plenty of labs have released model cards. Almost nobody has released the operational log of a frontier run. That is what makes this paper worth reading slowly.

The report covers the choice of model size, the data mix, the parallelism strategy, the failure record of the cluster, the annealing procedure, the context extension, and the post-training loop. Each of those is a place where labs usually say nothing, and each is a place where the rest of the field has been guessing.

Picking 405B from a scaling law

The compute budget was 3.8 times 10 to the 25 FLOPs, about 50 times Llama 2. Rather than trust a Chinchilla-style rule for how to split that between parameters and tokens, the team built its own scaling laws on runs from 6 times 10 to the 18 to 10 to the 22 FLOPs. The procedure has two stages. First, fit the relationship between compute and negative log-likelihood on downstream tasks. Second, fit the relationship between that log-likelihood and task accuracy. Chained together, the two fits predicted benchmark performance across four orders of magnitude of compute.

This is the part we would hand to anyone planning a large run. The size decision was made from predicted downstream accuracy rather than from a loss curve alone, and the prediction was checked afterwards. A team without Meta's compute can still copy the method at its own scale.

The failure log

The number people quoted most is 466. Over a 54-day snapshot of pretraining there were 466 job interruptions, 47 planned and 419 unexpected. GPU problems caused 58.7 percent of the unexpected ones. Despite that, effective training time stayed above 90 percent, which the paper credits to automated recovery and diagnosis tooling rather than to hardware reliability.

The system ran 4D parallelism, combining tensor, pipeline, context, and data parallelism, and achieved 38 to 43 percent model FLOPs utilisation on H100s with 80 GB of HBM3. Those two numbers together, the interruption rate and the utilisation, are the closest thing we have to a public baseline for what a frontier run costs in engineering rather than in dollars. Before this paper, the answer to how often a 16,000 GPU job fails was folklore.

Data mix and annealing

The pretraining mix was roughly 50 percent general knowledge, 25 percent mathematical and reasoning tokens, 17 percent code, and 8 percent multilingual. Context was extended from 8K to 128K in stages of continued pretraining rather than in one jump. And near the end of training the team annealed on a small amount of high-quality data, decaying the learning rate while upweighting the best sources.

The annealing result is the one we find most useful because it comes with a negative. On the 8B model, annealing on high-quality data improved GSM8K by 24.0 percent and MATH by 6.4 percent. On the 405B model the benefit diminished. So a trick that looks like free accuracy at small scale buys much less at the frontier, and the paper says so rather than reporting only the number that looked good.

Post-training as a loop

Post-training ran for six rounds, each combining supervised finetuning, rejection sampling, and direct preference optimisation. The loop shape is the finding. Rather than one SFT stage followed by one preference stage, the team repeatedly generated candidates from the current model, filtered them, trained on the survivors, and applied DPO, then started again from the improved model. The synthetic data that makes this work is one reason the license was changed to allow using Llama outputs to improve other models.

The blog post also describes quantising the 405B model from BF16 to FP8 for inference, which is how a model this size becomes deployable on a single node. The choice of a dense architecture over mixture of experts is explained as a stability decision.

What the field took from it

Looking back, the paper's value was less any single result than the fact of a complete account. Open training projects now had a reference for what a data mix, a parallelism layout, a failure rate, and a post-training loop look like at a scale they could not reach, and could calibrate their own choices against it. The scaling-law procedure and the annealing negative result are the two pieces we have seen cited most in practice.

What we would still want is the same level of detail from a second lab, so that we know which of these numbers are properties of frontier training and which are properties of Meta. One log is a data point. Two would be a field.

Sources

  1. Llama Team, The Llama 3 Herd of Models (arXiv 2407.21783)
  2. Meta AI, Introducing Llama 3.1: Our most capable models to date