State of open models, summer 2026: likes versus downloads
Hugging Face's summer report has Chinese labs shipping the largest open model nearly every month, models under 1B taking 83 percent of downloads, and coding agents overtaking humans as Hub users. Reading notes on which of those figures measure openness and which measure something else.
The report
Hugging Face published its summer 2026 State of Open Models report on 14 August, with a long author list led by Clement Delangue and about 150 contributors. It is built on Hub telemetry, which makes it the closest thing we have to a census of open model activity. The Hub grew from 2.43 million to 2.96 million model repositories between January and August, and from 711,000 to a million datasets. We have spent a week with it and want to sort its headline numbers into three bins. Numbers about who is releasing what. Numbers about who is using what. And numbers that only look like they are about openness.
Who is releasing the largest models
The release side is unambiguous. In almost every month of 2026 the largest and best performing open model from a Chinese lab was larger than any model an American lab released. The monthly ceiling from Chinese labs ran from 754B to 2.78 trillion parameters, with Kimi-K3 at 2.8T and Qwen releases in the 2.4T range. American releases stayed under 130B in five of seven months, with the exceptions being NVIDIA's Nemotron 3 Ultra at 561B and Thinking Machines Lab's Inkling.
Licensing follows the same split. Of 178 Chinese releases above 20B parameters, 59 percent carry Apache 2.0 and 22 percent MIT. American releases at the same size are 29 percent Apache or MIT and 41 percent custom terms. The report's own reading is that whatever these releases are for, it is not licence revenue, since the weights are given away on the most permissive terms available, and the value is captured through API and cloud business, hardware, or ecosystem position. That is the most direct statement of the economics of open weights we have seen from Hugging Face, and we think it is right.
These are openness numbers. Parameter count, licence, and cadence describe what a lab put into the commons and on what terms. On that measure, the centre of gravity of open weights moved to China some time ago and has not moved back.
Who is using what
The usage side tells a story that has nothing to do with frontier size. Models under 1B parameters account for 83 percent of all time downloads. Models above 100B account for 1 percent. In 2026 downloads above 70B are 3 percent of volume. In the GGUF format that local inference tools use, Qwen leads at 39.6 million monthly downloads, Gemma at 20.8 million and Llama at 7.5 million. The Qwen ecosystem has 151,448 derivative models on the Hub, 2.6 times Meta's total footprint and 4.7 times the Llama repositories specifically, and it is growing at 180 to 210 new repositories a day.
The section that gives the report its subtitle is the comparison of likes and downloads. Exactly one repository appears in both the top 25 by likes and the top 25 by downloads. all-MiniLM-L6-v2, a small sentence embedding model from 2022, has 1.55 billion downloads and 5,156 likes. Kimi-K3 gets about 60 downloads per like. Models from 2022 dominate the download rankings and no 2026 release makes the top 25. The report's summary is that likes are the instrument for reading what the field is excited about and downloads are the instrument for reading what it currently depends on.
We would go one step further. Downloads measure infrastructure. A model that is pulled 1.55 billion times is a dependency in a lot of production pipelines, and its openness matters in the boring sense that people would be in trouble if it went away. Likes measure attention. Attention is worth tracking, but it is not what the word open is for.
Agents as users
The third bin is the one we did not expect. For the first time, agents rather than people are the primary Hub users. The agent usage dataset published in July shows Claude Code with 44.4 percent of agent tagged traffic that month, down from 67.8 percent in April and 64 percent in May. Codex climbed from 10.4 percent to 20.8 percent over the same period. Nearly a quarter of July's agent traffic came from unnamed harnesses, and in May 59.8 percent came from unregistered agents. More than a dozen new client identifiers appeared between April and July. The report calls it a market with no incumbent, where one changed default can move half the traffic in a month. It also records the first documented autonomous agent intrusion in July, the point at which, in its phrase, an agent stopped being a reader and became an intruder.
We cannot defend calling this an openness number. It describes the interface through which people reach open artefacts. But it changes what the download figures mean. If most Hub reads are now made by coding agents fetching whatever their harness defaults to, then download counts are partly a measurement of harness defaults, and the 83 percent for sub-1B models is partly a fact about what agents pull for embeddings and tokenisers rather than what humans choose.
What we would want in the winter report
The report already separates likes from downloads. The next cut it should make is human from agent downloads, per model, so that the dependency figures can be read cleanly. We would also like to see the licence breakdown weighted by downloads rather than by release count, because the question that matters for the commons is not how many permissively licensed models were released but how much of the actual traffic runs on them. Our guess is that the answer is most of it, since Qwen leads GGUF downloads and Qwen releases are largely Apache 2.0. If that guess holds, then on the measure that counts, the open ecosystem in 2026 is more open than the release lists alone suggest, and the country that made it so is not the one most of the industry press is writing about.
Sources
From the foundation