The weekend LLaMA ran on a MacBook: llama.cpp and the 4-bit moment
Between March 10 and March 13, Georgi Gerganov's llama.cpp went from initial release to running a 7B model on a MacBook, a Raspberry Pi and a Pixel 6. An argument that this week, rather than any model release, is what created the local inference ecosystem.
Four days
On Friday March 10, Georgi Gerganov, a developer in Sofia, pushed the initial release of llama.cpp, a port of Meta's LLaMA inference code to plain C and C++ with no dependencies. The stated goal in the README was to run the model using 4-bit quantization on a MacBook. On Saturday Simon Willison wrote that he had it running, and that the 7B model occupied about 4 GB on disk with the 13B model just under 8 GB. On Sunday someone had it generating on a Raspberry Pi 4 with 4 GB of RAM, at roughly ten seconds per token. On Monday it ran on a Pixel 6, at 26 seconds per token.
None of those last two are useful speeds. That is not the point. Two weeks earlier LLaMA was a research release available to approved institutions, and the conventional wisdom was that a model of this class needed a datacentre GPU to run at all. By Monday evening the question had shifted from whether a language model could run on the device in your pocket to how fast.
What made it possible
Two things had to be true. The weights had to be available, and they were, because within days of the research release a pull request appeared on Meta's own repository linking to an unofficial BitTorrent download. And the memory footprint had to shrink by a factor of four, because a 7B model in 16-bit floating point is 13 GB, which does not fit in a MacBook's unified memory alongside anything else, and certainly does not fit in a Raspberry Pi.
The shrinking is where llama.cpp earned its place. Storing each weight as a 4-bit integer instead of a 16-bit float cuts the file to a quarter and, more to the point, cuts the memory bandwidth needed per token by the same factor. On a laptop, bandwidth is the constraint, so a 4-bit model is both smaller and faster to run. The quality loss from crude 4-bit rounding is real and measurable, and nobody in that first week had measured it carefully. What they had shown was that the loss was small enough that the model was still obviously a language model.
The other ingredient was the code itself. Gerganov had spent the preceding months building ggml, a tensor library he started in September 2022, and had already used it to port OpenAI's Whisper speech recognition model to C++ as whisper.cpp. So llama.cpp was not written in a weekend. The weekend was the last step of a six-month project that happened to be ready the week the weights leaked, and it used the ARM NEON and Accelerate paths on Apple silicon that the earlier project had already exercised.
Why this and not the model
The argument we want to make is that llama.cpp, rather than LLaMA, created the local inference ecosystem. LLaMA was necessary. It was the first strong model whose weights were, however unintentionally, in public hands. But a set of weights that runs only on hardware most people do not own is a paper, in effect, with a very large appendix. What turns weights into an ecosystem is the ability of a person with a laptop to run them, modify them, and build on top of them without asking anyone.
Willison's phrase for this was that language models were having their Stable Diffusion moment, and the analogy is exact in one respect. Stable Diffusion's release in August 2022 mattered because it ran on a consumer GPU, and the tooling, fine-tunes and interfaces that followed were built by people who could iterate on their own machines. The models that were better but hosted did not generate that activity. Capability that you cannot run locally produces users. Capability that you can produces developers.
There is a second reason the tool matters more than the model. Models get replaced. Whatever succeeds LLaMA will need to be quantised and run on a laptop too, and the code path for doing that now exists, is written in a language every platform compiles, and is being extended by dozens of contributors in its first week. The model is a snapshot. The runtime is infrastructure.
What is still missing
A lot, and we want to be specific rather than triumphant. Nobody has published a careful quality comparison of the 4-bit LLaMA against the 16-bit original on a standard benchmark. The 4-bit scheme in use is the simplest possible one, and there is a literature on post-training quantisation that has not yet been applied. Prompt formats, sampling settings and the tokenizer are all being reverse-engineered from the original code, and the chat-style behaviour people expect from ChatGPT is absent because LLaMA is a base model with no instruction tuning.
Legally the situation is unresolved. The weights are under a non-commercial research licence that most of the people running them did not agree to, and Meta has not said what it intends to do about the torrent. That will not stop anyone this week. It may shape what companies are willing to build on this in the months ahead.
What we would try
The obvious next experiment is the quality measurement. Take the 7B and 13B models, run them at 16-bit and 4-bit on perplexity over a fixed held-out text, and report the gap. If the gap is small, the case for 4-bit as a default is made. If it is not, the next step is applying the better quantisation methods from the literature, which trade a little more compute at conversion time for less damage to the weights.
The less obvious experiment is instruction tuning on a laptop. If inference fits on consumer hardware, low-rank fine-tuning might too, and a 7B base model plus a few thousand instruction examples is a very different thing from a 7B base model alone. We would expect someone to try this within weeks, because the tooling from this weekend makes the attempt cheap, and cheap attempts are what an ecosystem is made of.
Sources
From the foundation