What actually happened

On February 24 Meta announced LLaMA, a family of four models at 7, 13, 33 and 65 billion parameters, trained on 1.4 trillion tokens of public data including Common Crawl, GitHub, Wikipedia, Project Gutenberg, arXiv and Stack Exchange. Access was by application, limited to researchers at institutions, government agencies and NGOs, under a noncommercial licence. Meta's paper reported that the models beat GPT-3 on all twenty zero and few-shot tasks it tested, and that the 13B model outperformed the 175B GPT-3 on most benchmarks.

Around March 1, roughly a week after the announcement, someone posted a BitTorrent link to the full weights on 4chan. Vice reported it on March 7 as the first time a major tech company's proprietary language model had leaked to the public. Copies went up on GitHub and Hugging Face within days. Meta filed takedown requests, and at least one GitHub repository was flagged by Meta for unauthorized distribution and taken offline. Meta's public statement was measured: some people had tried to circumvent the approval process, but the company believed its release strategy still balanced responsibility and openness.

The takedowns did not matter. A torrent cannot be recalled, and by the time the notices went out the weights were on thousands of machines. What matters for the rest of this post is what people did with them in the next three weeks.

Three weeks of building

On March 10, Georgi Gerganov released llama.cpp, a C++ inference engine whose stated goal was to run the model with 4-bit quantization on a MacBook. The 7B model fit in about 4GB on disk and the 13B in about 8GB. Simon Willison ran the 13B model on his laptop the next day and wrote that language models were having their Stable Diffusion moment, meaning the point at which a capable model becomes something you run yourself rather than something you rent through an API. Within days people reported the 7B model running on a Raspberry Pi 4 with 4GB of RAM at around ten seconds per token, and on a Pixel 6 phone at 26 seconds per token.

On March 13, Stanford's Center for Research on Foundation Models released Alpaca, a 7B LLaMA fine-tuned on 52,000 instruction-following examples generated with OpenAI's text-davinci-003. The data cost under $500 in API calls and the fine-tune under $100 on cloud compute, about three hours on eight A100s. In a blind human comparison against text-davinci-003 Alpaca won 90 of 179 pairings, roughly even. Stanford released the data, the generation code and the fine-tuning code, and said the weights would follow pending Meta's approval.

Then Vicuna, from a student team across Berkeley, CMU, Stanford, UC San Diego and MBZUAI, trained a 13B LLaMA on around 70,000 user-shared ChatGPT conversations from ShareGPT for about $300. Their headline claim, that Vicuna reached over 90% of ChatGPT's quality, came with a footnote calling the evaluation fun and non-scientific because GPT-4 was the judge. We would not repeat the number without the footnote. The point is the cost and the timeline, both of which were unthinkable a month earlier.

Why the leak, and not a policy, set the template

It is worth being clear about what Meta's actual policy was. It was a research licence with case-by-case approval. That is a reasonable policy, and it would have produced a modest amount of academic work over a year or two. It would not have produced llama.cpp, because Gerganov was a solo developer in Sofia, not an applicant from an approved institution. It would not have produced Alpaca as a public release, because every derivative would have inherited the same gated access.

What the leak did was remove the approval step from the loop. Once the weights were simply available, the binding constraint shifted from permission to compute, and the community immediately showed that the compute constraint was much softer than assumed. Quantization got a 65B model onto a single A100 and a 13B onto a MacBook Pro. Fine-tuning a useful instruction model cost less than a conference registration.

The licences did not change. Alpaca and Vicuna both remain noncommercial, bound by LLaMA's terms and by OpenAI's terms on the data used to build them. So the leak did not create a legal open-source model. What it created was a demonstration that a lab can release weights to the world and the result is a wave of tooling and derivatives rather than a catastrophe, at least so far. That demonstration is now on the record, and every lab deciding how to release its next model will be arguing against it.

What we are not sure about

Willison's post lists the misuse cases plainly: spam, romance scams, trolling, disinformation, automated radicalization. None of those need a 65B model, and most were possible with GPT-3 through the API, but a local model removes the last point of control. We do not think three weeks is enough time to know whether that cost is large. We also do not think anyone should pretend the 4chan release was a considered decision by anyone. It was one person with a torrent client.

The benchmarks are also weaker than they look from the outside. LLaMA's numbers come from Meta's own paper, and the derivative models are being evaluated by other language models or by small human panels. We do not yet have an independent evaluation of any of this, and we would treat every quality claim from March 2023 as a hypothesis until someone outside the releasing team reproduces it.

What we would want to see next

The obvious experiment is a deliberate open release with a permissive licence, so that the next Alpaca is something a company can ship. If the LLaMA leak is the accidental version of the open-weights era, someone will build the intentional version, and we want to see whether the derivative ecosystem looks the same when nobody has to worry about the licence.

For our own work, the practical change is that we can now run a model of this class on lab hardware, inspect its activations, and fine-tune it on data we control. That was not true in February. Whatever else the leak was, it gave every independent research group a subject to study that is not behind someone else's API, and we expect the next year of interpretability and evaluation work to look different because of it.

Sources

  1. The Batch: How Meta's LLaMA model leaked
  2. Vice: Facebook's powerful large language model leaks online
  3. Simon Willison: Large language models are having their Stable Diffusion moment
  4. Stanford CRFM: Alpaca
  5. LMSYS: Vicuna