What open source AI means now: OSAID, Llama and OLMo 2
OSI published a definition of open source AI in October, Meta said the bar was too narrow, the Software Freedom Conservancy said it was too low, and then Ai2 shipped OLMo 2 with everything on the table. The argument is now about a real model.
What the definition requires
The Open Source Initiative published version 1.0 of the Open Source AI Definition on October 28 at All Things Open in Raleigh, after a multi-year drafting process. It applies the four freedoms of open source software, use, study, modify and share, to AI systems, and it names the artifacts a release must include to grant them. Those are the source code used to train and run the system under an OSI-approved licence, the model parameters under terms that permit free use, study, modification and redistribution, and what the definition calls data information.
Data information is the contested clause. OSI does not require the training dataset itself. It requires enough detail about the data, its provenance, its processing and how to obtain or license it that a skilled person could substantially recreate the system. OSI's reasoning is that much of the data behind current models is encumbered by copyright, contract or privacy law, and that a definition requiring the full dataset would exclude nearly every real system and leave the term to vendors who release nothing.
Under this test Llama fails on two counts. Its licence restricts commercial use by services with more than 700 million users and prohibits certain uses outright, and Meta does not disclose the training data in enough detail to recreate it. Stable Diffusion and Mistral's Ministral models fail on licence terms too. So the first thing the definition did was tell the most-downloaded open-weight models that they are not open source.
Too narrow, or too low
Meta's response was that there is no single open source AI definition and that earlier definitions do not encompass the complexity of today's models. Meta disagreed with the pronouncement while saying it agreed with OSI on other matters. Read plainly, that is a company that wants to keep calling Llama open source and does not accept that a standards body gets to say otherwise.
From the other side, Bradley Kuhn at the Software Freedom Conservancy wrote on October 31 that the definition erodes the meaning of open source, because it fails to require reproducibility by the public of the scientific process of building these systems. His position is that the ideal is to train only on publicly available data under free licences, though he is candid that he does not yet know for certain whether that is the only way to respect user rights. His proposal was that OSI rebrand 1.0 as current recommendations rather than a definition, and he said he would run for the OSI board on that basis.
We find ourselves closer to Kuhn on principle and closer to OSI on tactics. A definition that excludes every model that exists is a manifesto, and a definition that includes Llama is marketing. Data information is a compromise that lets a small number of systems qualify today while making clear what the rest are missing. Whether it is the right compromise depends on whether anyone can actually meet it at competitive quality. Four weeks later someone did.
OLMo 2 as an existence proof
On November 26 Ai2 released OLMo 2 at 7B and 13B parameters, trained on up to 5 trillion tokens, with weights, training data, code, recipes, intermediate checkpoints and instruction-tuned variants all public. Stage one trained on OLMo-Mix-1124, about 3.9 trillion tokens drawn from DCLM, Dolma, Starcoder and Proof Pile II. Stage two used Dolmino-Mix-1124, 843 billion tokens split evenly between filtered web data and domain-specific high-quality content, sampled into 50B, 100B and 300B token mixes. Post-training followed the Tülu 3 recipe of supervised fine-tuning, DPO and reinforcement learning with verifiable rewards.
The benchmark claims are that OLMo 2 7B outperforms Llama 3.1 8B and that OLMo 2 13B outperforms Qwen 2.5 7B with fewer training FLOPs, on Ai2's own OLMES suite of 20 benchmarks. The 13B instruct model is described as competitive with Qwen 2.5 14B Instruct and Llama 3.1 8B Instruct. These are the releasing lab's numbers on the releasing lab's evaluation suite, and we would want an outside replication before repeating them as fact. But the ordering is plausible and the gap to Llama is small enough that data information is no longer a claim about hypothetical systems.
What OLMo 2 does not resolve is Kuhn's objection. The training mix is documented and downloadable, but it includes web crawl whose individual documents carry no licence. It satisfies OSAID because a skilled person can recreate it. It does not satisfy the stricter standard where every input is freely licensed, and we do not know of a competitive model that does.
What the argument is now about
Before October the question was whether open source AI meant anything. Now there are three concrete positions. Meta's is that weights plus a restricted licence is open enough. OSI's is that weights, code and a recreatable data description are the floor. The Conservancy's is that the data itself must be free. OLMo 2 sits on OSI's line and shows it can be held at 7B and 13B without giving up much.
The open question for us is whether it can be held at frontier scale. The larger the model, the more the training set depends on data nobody can fully describe, and the more expensive it gets to publish intermediate checkpoints. If the line only holds below some size, then open source AI will be a category for small models, and everyone will know it.
What we would like to see is the same comparison at 70B, and we would like the Conservancy's standard tested empirically too. Someone should train on nothing but freely licensed data at a size where it matters and report how far behind it lands. If the answer is a few points, that changes the argument. If the answer is a generation, that changes it differently.
Sources
From the foundation