The ratio everyone had agreed on

Since last spring the working rule for training a language model has come from Hoffmann et al. at DeepMind, the Chinchilla paper. Given a fixed training budget, you get the lowest loss by growing parameters and tokens together, roughly twenty tokens per parameter. By that rule a 10B model should see about 200B tokens. Train it longer and you would have been better off with a bigger model instead.

Meta's LLaMA paper, out this week from Hugo Touvron, Guillaume Lample and a team at FAIR, trains a 7B model on 1.0T tokens. The 13B model also sees 1.0T, and the 33B and 65B models see 1.4T. The 7B run is roughly seven times past the Chinchilla ratio. This was a choice, and the introduction says exactly why.

Training compute is the wrong budget

The argument is that the Chinchilla objective answers the wrong question for most users of a model. It minimises loss for a given training budget. But a model is trained once and then served many times, and the cost that dominates over its life is inference. For a target level of quality, the model you want is the smallest one that reaches it, because that is the one that is cheapest to run. The way to make a small model reach a given quality is to train it on more tokens than the compute-optimal recipe would allow.

The paper reports that the 7B model's performance was still improving when training stopped at one trillion tokens. That single observation carries the argument. If the curve had flattened at 200B tokens, Chinchilla would have been right about where to stop and Meta would have wasted compute. It did not flatten. So the extra tokens bought real capability in a model that fits on one GPU.

There is a cost, and the paper is open about it. The 65B run used 2,048 A100 80GB GPUs for about 21 days, and the total across the family came to roughly 1,022,362 GPU hours. Meta is paying more at training time to hand out models that are cheaper at serving time. For a lab whose stated aim is to release the weights to researchers, that trade is the right one.

What the small models can do

The benchmarks make the case concrete. LLaMA-13B outperforms GPT-3 at 175B on most of the common sense reasoning tasks the paper reports, with a model about a tenth the size. On NaturalQuestions in the 64-shot setting it reaches 31.9 percent exact match against 29.9 percent for GPT-3. At the top of the family, LLaMA-65B is competitive with Chinchilla-70B and with PaLM-540B on the question answering tasks.

A 13B model that matches a 175B model is a different object from a research point of view. It runs on a single accelerator with modest memory. It can be fine-tuned by a university group. The comparison to PaLM is the headline, but the comparison to GPT-3 is the one that changes who can do the work.

Public data only

The second claim of the paper is that all of this was done with publicly available data. The mixture is 67 percent CommonCrawl, 15 percent C4, 4.5 percent each of GitHub, Wikipedia and books, 2.5 percent arXiv and 2 percent Stack Exchange. No proprietary corpus, no licensed data. The authors say this explicitly as a point of principle, that state-of-the-art models can be trained without access to datasets nobody else can inspect.

That matters for reproducibility more than for performance. If the data recipe is public, a result is checkable, and a failure mode can be traced to a source. It also means the overtraining argument is testable by anyone with the compute, because the tokens are there to be re-collected.

What this changes

We expect the token-to-parameter ratio to stop being read as a law and start being read as a training-only optimum. Chinchilla told us how to spend a fixed training budget to minimise loss. It never claimed to tell us what model to build for deployment, and the field had been treating it as if it did. LLaMA reframes the question as picking a size for the inference budget and then training past the point where the compute-optimal recipe stops.

The experiment we would like to see is the one the paper implies but does not run. Take the 7B model past one trillion tokens, to two or three, and report where the curve finally flattens. If it keeps going, the current ratio for small open models is still too low. If it stops, we will have a number for the ceiling on overtraining, which is the thing every group planning a small model now needs.

Sources

  1. Touvron et al., LLaMA: Open and Efficient Foundation Language Models (arXiv 2302.13971)
  2. LLaMA, full text (ar5iv)