What the document is

SemiAnalysis published the memo on May 4. It was shared by an anonymous person on a public Discord server and is attributed to a researcher at Google. Nobody has confirmed the authorship, and Simon Willison, whose write-up is where most of us read it, makes the sensible point that the argument stands or falls on its own regardless of who typed it. We are going to treat it the same way.

The claim is simple. Google and OpenAI are both being outrun by open models, and neither has a durable advantage. The evidence offered is the two months since Meta's LLaMA weights got out: Alpaca showed instruction tuning could be done for a few hundred dollars, Vicuna arrived a week later with better quality, and after that came quantised builds, RLHF variants and multimodal bolt-ons, each within days of the last. The line everyone is quoting is that people are doing things with 100 dollars and 13 billion parameters that Google struggles with at 10 million dollars and 540 billion.

The technical argument is mostly about LoRA

Strip the rhetoric away and the memo is a bet on one technique. Low-rank adaptation lets you fine-tune a large model by training a small set of extra matrices, which means the job fits on consumer hardware and finishes in hours. The memo's author adds a second observation that we think is the real insight. LoRA updates are small and stackable. You can train one adapter for a task, someone else trains another, and they compose. Improvements accumulate across a community without anyone needing to retrain the base model.

The corollary is that a lab which starts a large training run from scratch every few months is throwing away the accumulated adaptations of the previous version. The memo argues this makes giant training runs a liability rather than a moat, because the open ecosystem iterates in weeks on top of a fixed base while the lab spends months rebuilding the base.

From where we sit, the LoRA part is right and we have been living it. We do not have a 540 billion parameter model and we never will. We do have a handful of GPUs and a queue of adapters, and the speed at which a small group can now go from an idea to a fine-tuned model that tests it is unlike anything we have worked with before.

Where we think the memo overreaches

The memo slides from 'the open ecosystem iterates faster' to 'the gap is closing astonishingly quickly' and treats those as the same claim. They are not the same claim. Fast iteration on 7 and 13 billion parameter models is real. Matching the best closed model is a different thing, and none of the timeline entries the memo lists is evidence of it. Vicuna's quality claim came from asking GPT-4 to grade it against ChatGPT, which is a comparison that rewards a model for sounding like the judge.

There is also a base rate problem in the memo's own story. Every item in the timeline depends on LLaMA existing. The hobbyists did not train a foundation model for 100 dollars. Meta trained one for a great deal more and the weights leaked. The 100 dollar figure is the cost of the last step, and the memo's framing invites the reader to forget the first one. If Meta or someone like them stops releasing bases, the ecosystem the memo describes has nothing to fine-tune.

The memo also underrates what money buys at the top. The claim that the quality gap is closing rests on model-graded comparisons of chat quality. Evaluating whether a 13B model with an adapter matches a frontier model on hard reasoning, long context or reliability under distribution shift is work that the memo does not do, and the people in the timeline have not done it either.

What it gets right about institutions

The part we found most persuasive has nothing to do with parameters. The memo argues that Google's secrecy buys little, because researchers leave and take what they know with them, and that the lab would be better off learning from the open community than trying to wall it out. Willison draws the parallel to what happened when Stable Diffusion's weights were public and a whole tooling ecosystem grew up in months that no closed image model had.

That argument is about where ideas come from, and we think it is correct regardless of how the capability race goes. The interesting techniques of the last two months, instruction tuning at low cost, quantisation to run on a laptop, adapters that stack, all came from outside the two labs the memo is addressed to. A lab that treats that stream as a threat rather than an input is making a choice about its own research diet.

Reading it from a small lab

For a group our size the memo is less a prophecy than a description of what became possible in March. We can now run a meaningful experiment on a real language model with a single machine, which last year would have meant begging for API credits and accepting that the model was a black box. That is the durable change and it is worth being clear about it separately from the moat question.

On the moat question itself we would take the other side of the memo's strongest claim. Fast iteration on small models is here to stay. Frontier parity from that iteration is a prediction, and we would want to see a 13B model with a stack of adapters beat GPT-4 on an evaluation GPT-4 did not grade before we believed it. We would be happy to be wrong, and if the memo's author is right we will know within the year.

Sources

  1. Simon Willison, Leaked Google document: "We Have No Moat, And Neither Does OpenAI" (May 4, 2023)