Mistral goes back to Apache 2.0: what a license reversal signals
Mistral Small 3 shipped under Apache 2.0 with a public promise to move general-purpose models away from the research-only MRL license. A look at what the restrictive license was buying, what it was costing, and why the reversal landed the week it did.
What shipped and under what terms
Mistral Small 3 came out yesterday, January 30, as a 24 billion parameter model under Apache 2.0. The announcement puts it at over 81 percent on MMLU, around 150 tokens per second on their serving stack, and roughly three times faster than Llama 3.3 70B instruct on the same hardware. The model was trained without reinforcement learning and without synthetic data, which the company frames as a clean base for people who want to do their own post-training. Quantised, it fits on a single RTX 4090 or a 32 GB MacBook.
The benchmark story is interesting but the line that matters most sits near the bottom of the post. Mistral says it is renewing its commitment to Apache 2.0 for general purpose models and progressively moving away from MRL-licensed models. That is a reversal, stated plainly, and it is worth pausing on because licenses rarely move in the permissive direction once a company has tightened them.
What the MRL was
The Mistral AI Research License arrived with the September 2024 releases. The previous Mistral Small, the 22 billion parameter v24.09, shipped under it, while Pixtral 12B from the same day went out under Apache 2.0. The MRL grants a royalty-free, worldwide right to use, modify, and distribute the weights, but only for research purposes, which the text defines as personal, scientific, or academic work that is non-profit and non-commercial.
The prohibitions are broad. Employees using the model in a workplace context count as commercial. So does any activity intended to generate revenue, any distribution by a commercial entity even free of charge, and any offering through SaaS or cloud instances. Derivatives can be shared under different terms as long as those terms do not contradict the original, and every redistribution has to carry the license text and a Mistral attribution notice. In practice this meant a startup could evaluate the model and then had to call sales.
What a restrictive license protects and what it costs
The logic of the MRL was clear enough. Give researchers the weights, keep the commercial deployment revenue, and stop cloud providers from hosting the model without a deal. That is a reasonable bet if your model is the one people want and the alternatives are worse. It stops being reasonable the moment a comparable Apache or Llama-licensed model exists, because at that point the license is the only thing separating your download from theirs, and it separates them in the wrong direction.
The costs are mostly invisible from inside a company and very visible from outside it. Fine-tuners pick the base they can ship. Inference providers integrate the models they can host. Tutorial authors, benchmark maintainers, and quantisation projects follow the models with the fewest legal questions. Each of those choices is small, and together they decide which model becomes the default that everything else is measured against. Mistral 7B and Mixtral earned that position in 2023 under Apache 2.0. The MRL models did not, and the gap in community activity was the signal.
The Small 3 post reads like a company that did the arithmetic. It positions the model as an open replacement for closed models like GPT-4o mini, and lists Hugging Face, Ollama, Kaggle, Together AI, Fireworks AI, and IBM watsonx as launch partners. None of that distribution works under a license that forbids hosting. The reversal had to happen before the launch could.
The week it landed
The timing is hard to separate from the week's other news. DeepSeek-R1 arrived ten days earlier under an MIT license, with a full technical report, and pulled attention toward permissive weights from a lab nobody in Paris or San Francisco controls. A European lab announcing a research-only license in that environment would have looked like it was fencing off a field everyone else had just opened. Whatever the internal timeline for the Apache decision, the public statement now reads as a response to a market that had moved.
We do not think Mistral was wrong to try the MRL. Building frontier-adjacent models is expensive and someone has to pay for it. What the experiment showed is that a mid-sized open model cannot charge for the license, because the license is the product feature that makes people choose it. Revenue has to come from somewhere else, from hosted APIs, from enterprise deployments, from fine-tuning services, and the weights are the marketing budget.
What to watch
The promise is progressive, and progressive promises deserve auditing. The obvious test is whether the next Mistral Large or the next Codestral ships under Apache 2.0 or stays behind the MRL with a note that it is a specialised model rather than a general purpose one. The wording leaves that door open. We would also like to see whether the community response to Small 3 looks like the response to Mistral 7B, with fine-tunes and quantisations appearing within days, because that would confirm the license was the thing holding the previous models back.
For those of us who train on and study these models, the practical change is that a 24B Apache base with no RL and no synthetic data is now available. That is a good object for post-training experiments precisely because it has not been pushed toward anyone's preference distribution yet. We would rather run a DPO ablation on that than on a model that has already been through three rounds of someone else's alignment.
Sources
From the foundation