What changed on 18 July

Three weeks ago Meta released Llama 2, and in our group the effect was immediate. Every fine-tuning script we had been running on the leaked first-generation weights got pointed at the new checkpoints within a day, and for the first time the results were something we could put in front of a partner without a conversation about licensing. The models come in 7B, 13B and 70B sizes, pretrained on 2 trillion tokens, 40 percent more than Llama 1, with a context window doubled to 4,096 tokens. A 34B model was trained but held back pending further red-teaming. The 70B uses grouped-query attention, which makes serving the large model cheaper on memory.

The two things that matter most for what happens next are not the base weights. They are the chat variant, Llama 2-Chat, which ships alongside each size, and the licence, which permits commercial use. The base weights of Llama 1 were already in circulation. What nobody had was a well-aligned chat model you were allowed to build a product on.

Why the licence is the release

The Llama 2 Community License grants a worldwide, royalty-free licence to use, modify and distribute the models, including commercially. It has two conditions worth knowing. If your products exceed 700 million monthly active users you must request a separate licence from Meta, which Meta may grant at its discretion. And you may not use the models or their outputs to improve any other large language model outside the Llama family. Distributed copies must carry an attribution line.

The 700 million clause excludes a handful of companies and nobody else. The output clause is the one that will generate argument, since it means you cannot legally distil Llama 2 into a competitor, though you can distil it into another Llama. For a startup, a university group or a foundation like ours, neither restriction bites. That is why the ecosystem is forming around this release rather than around any of the other open models of the past year. Most of those were research-only, and a research-only licence is a promise to lawyers that you will never ship.

What the paper reveals about RLHF at scale

The paper is unusually detailed about the fine-tuning pipeline, and that detail is the reason it will be read for years. Supervised fine-tuning used 27,540 annotations, a small number that the authors argue was enough because the quality was high. Preference data was much larger, around 1.4 million binary comparisons collected in-house and combined with open datasets from Anthropic and Stanford. Two reward models were trained rather than one, a helpfulness model and a safety model, because the authors found the two objectives pulled against each other and a single model could not serve both.

RLHF ran in five iterations, V1 through V5. Early rounds used rejection sampling, generating several candidates and fine-tuning on the reward model's favourite. Later rounds added proximal policy optimisation on top. The paper also introduces Ghost Attention, a training trick that keeps a system instruction in force across many turns of a conversation by synthetically attaching it to every turn during training and then dropping it at inference. Safety was handled with its own RLHF signal plus context distillation, where a safety preamble is used to generate good responses and then removed so the model learns the behaviour without the prompt.

Two lessons in here run against what we had assumed. The first is that supervised data is the cheap part and preference data is where the scale goes. Twenty-seven thousand examples versus 1.4 million comparisons is a ratio most groups had backwards. The second is that the reward model, not the policy, is the piece that determines the ceiling. The authors say as much, and it fits what we see in our own much smaller runs.

The ecosystem that is already forming

The tooling is arriving faster than we can track it. Low-rank adapters, LoRA and its quantised cousin QLoRA, mean that a 7B model can be fine-tuned on a single consumer GPU by training a few million parameters on top of frozen weights. Frameworks like Axolotl wrap that into a config file, so producing a domain-specific Llama 2 variant is now an afternoon of work rather than a research project. Hugging Face is filling up with adapters for legal text, medical notes, code and every language with a community large enough to collect data.

The chat variant is what makes the adapter approach worthwhile. Fine-tuning a base model into a helpful assistant requires the whole pipeline the paper describes. Fine-tuning an existing assistant to know about your domain requires a few thousand examples and an adapter. Meta paid for the expensive part once and released the result, and everyone else is now doing the cheap part in parallel.

The commercial licence closes the loop. An adapter trained on Llama 2-Chat can be shipped. That was never true for the earlier open models, and it is the difference between a hobbyist ecosystem and one that pays people.

What we are watching

Two things worry us about this ecosystem and both are consequences of its speed. The first is that adapters inherit the alignment of the base chat model and then overwrite bits of it. Nobody is re-running the safety evaluations after fine-tuning, and a few thousand domain examples can move the refusal behaviour a long way. The second is the output clause. Distillation from Llama 2 into other Llama 2 models is permitted, so the family will become a closed loop of models trained on each other's outputs, and we do not think anyone knows what that does to quality after a few generations.

The experiment we would like to see is a simple one. Take Llama 2-Chat 7B, fine-tune it on a benign domain with a standard adapter recipe, and run the paper's own safety evaluation before and after. If the safety score holds, the ecosystem is on firmer ground than we fear. If it drops, we should know that now, while the number of shipped adapters is still counted in hundreds.

Sources

  1. Llama 2: Open Foundation and Fine-Tuned Chat Models (arXiv 2307.09288)
  2. Llama 2 paper, full text
  3. Llama 2 Community License