llama.cpp joins Hugging Face: who owns local inference now?
Georgi Gerganov and the ggml team joined Hugging Face on February 20 with a promise of full technical autonomy and a project that stays open source and community driven. The same month Qwen 3.5 shipped an 807GB open model at the top of a family that runs down to 0.8B. On what it means when the tooling that makes open weights usable consolidates under one hub.
What happened
On February 20 Hugging Face announced that Georgi Gerganov and the ggml team behind llama.cpp are joining the company. The ggml.ai site now says plainly that the company, founded in 2023 with pre-seed funding from Nat Friedman and Daniel Gross, was acquired by Hugging Face in 2026. The announcement calls llama.cpp the fundamental building block for local inference, which is a fair description. If you have run an open-weight model on a laptop in the last three years, there is a good chance the code doing the arithmetic came from this project.
The terms, as stated, are generous to the project. Georgi and the team keep 100 percent of their time on llama.cpp and have full autonomy and leadership on technical direction. The project will remain 100 percent open source and community driven. Hugging Face's contribution is described as long-term sustainable resources. The library and its relatives stay under the MIT licence they have always had. Nothing in the post mentions a licence change, a governance body, or a commitment period, and we read the absence as meaningful.
Why this month
The timing puts the acquisition next to the largest open release of the year so far. Qwen 3.5 started shipping on February 17 with a 397B model that uses 17B active parameters per token and takes 807GB on disk, followed by 122B, 35B, 27B, 9B, 4B, 2B and 0.8B variants, all with reasoning and vision. Simon Willison's notes on the family report good feedback on the 27B and 35B for coding on 32GB and 64GB Macs, and find the 9B, 4B and 2B models effective for their size. The 2B is 4.57GB in full precision and 1.27GB quantised.
That last pair of numbers is the whole story of why llama.cpp matters. The open-weights ecosystem is not open in any useful sense unless there is a way to run the weights on hardware people own, and for most people that means a quantised GGUF file and the llama.cpp runtime. An 807GB release is a research artefact. A 1.27GB file that runs on a phone is a product. The conversion between them is the thing Georgi's team built, and it is the thing Hugging Face now employs.
The consolidation question
Hugging Face already hosts the weights. It already maintains transformers, the reference implementation most open models ship with. It runs the model cards, the leaderboards, the datasets, and a large fraction of the inference endpoints. Adding the dominant local runtime means that for an open model, the place you download it from, the code you load it with, the format you convert it to and the engine you run it on can all be maintained by one company. Each of those was built by different people for different reasons. Now most of them have the same employer.
We want to be careful here, because the objection is structural rather than about intent. Hugging Face has been a good steward of the things it hosts, and the ggml announcement reads as a rescue of a project that was carrying a great deal of the ecosystem on a small team with pre-seed money. Sustainable funding for infrastructure is good. The concern is that the community's alternatives get thinner each time this happens, and the promise of autonomy rests on the goodwill of the current management rather than on any structure that would survive a change of ownership, a change of strategy, or a bad year.
What autonomy would look like on paper
There are ways to make the promise checkable. The project could be moved into a foundation or a neutral organisation with a trademark and a governance document, the way large open-source projects have done for two decades. The GGUF format could get a written specification maintained outside any one company, so that other runtimes can implement it without tracking a moving target. The commit rights and release process could be documented so that the community knows who decides and how that changes if the team leaves. None of that is in the announcement, and none of it is incompatible with what the announcement promises.
The reason to ask now rather than later is that the incentives are aligned now. Hugging Face wants to be seen as a good home for the project, and the team wants its independence on the record. Two years in, when the runtime has been integrated into the hub's products and the format has become load-bearing for the company's business, the same conversation will be harder to have. Open science depends on tooling that stays open when the funding changes, and the time to write that down is before it is tested.
What we would want next
Three things, all cheap. A GGUF specification published as a standalone document with a version number. A statement of who holds the llama.cpp trademark and the repository organisation, and what happens to them if the team or the company changes course. And an honest count from the community of how many independent runtimes can load a GGUF file today without depending on ggml code, because that number is the real measure of how much of local inference one company now owns. Our guess is that it is smaller than most people assume, and that is a fine reason to start building the alternatives while the current arrangement is friendly.
Sources
From the foundation