A magnet link and a 7B model: Mistral's release as a cultural statement
Mistral 7B arrived under Apache 2.0, via a BitTorrent magnet link, with benchmarks over Llama 2 13B and no paper attached. What the release style says about a five-month-old European lab and about what counts as publishing a model.
What arrived
On September 27 Mistral AI released a 7.3 billion parameter model under the Apache 2.0 license, and the first way to get it was a BitTorrent magnet link. The Hugging Face upload and a tar archive followed, along with instructions for deploying on AWS, GCP, and Azure using vLLM and SkyPilot. There was no paper. There was a blog post with benchmark tables and an architecture note.
The company itself is five months old. It was founded in April by Arthur Mensch, Guillaume Lample, and Timothée Lacroix, and raised 105 million euros in a seed round in June at a valuation reported around 240 million euros. So the first product of one of the largest seed rounds in European tech is a small model, a permissive license, and a torrent.
The claims
The post says Mistral 7B outperforms Llama 2 13B on all benchmarks, approaches CodeLlama 7B on code, and is on par with Llama 34B across many evaluations. On reasoning tasks it is described as matching a Llama 2 model more than three times its size. An instruct-tuned variant is claimed to beat all other 7B models on MT-Bench, and the post says no proprietary data went into that fine-tune.
Two architectural choices are named. Grouped-query attention, for faster inference. And sliding window attention with a 4,096 token window, which the post says gives linear compute in sequence length, roughly doubles speed on 16k sequences with FlashAttention, and caps the cache at the window size, saving half the memory on 8,192 token sequences. Both are known techniques. What is new is a small model that combines them and beats a larger one.
We have not reproduced any of these numbers, and neither has anyone else two days in. The benchmark tables are the company's own. That is the cost of releasing without a paper, and it is worth holding the claims lightly until the community runs them.
What the torrent says
A magnet link is an unusual delivery mechanism for a funded company. It has no gate, no click-through license, no account, and no way to count downloads. Choosing it is a statement about who the weights are for. Compare the Llama 2 flow, which requires a form and a custom license with usage restrictions. Mistral's choice says the model belongs to whoever pulls it, on the same terms as any Apache project.
It is also a statement about where the lab sits in the ecosystem. Meta has spent 2023 arguing that open weights are the responsible path and has done so with a license that is open in spirit and closed in letter. A European lab releasing under Apache 2.0 and skipping the form is out-opening Meta, and doing it with a model small enough to run on a laptop. That makes a point about European capability and about who gets to define open.
What counts as publication
The part of this we keep turning over is the absence of a paper. For most of the history of this field, a model was published when a paper described it, and the weights were a courtesy. Mistral inverted that. The weights are the publication. The benchmark table is the abstract. If you want to know what the model does, you download it and run it.
There are things a paper gives you that a checkpoint cannot. Training data composition, token count, compute budget, the ablations that justify the design, the failure cases the authors saw. None of that is in the post. What we have instead is a verifiable artefact that anyone can probe, which is more than many papers offer. We do not think one replaces the other, and we expect a paper will follow. But the order says something about what the lab thinks matters first.
What we would like to see next
Independent evaluations, first. If a 7B model really sits with Llama 2 13B across the board, the cost of a capable local assistant just halved, and the fine-tuning community will move onto it within weeks. Then the training details, because a result nobody can reproduce is a product claim rather than a finding.
And we would like the release style to be copied, the license and the lack of a gate more than the torrent itself. If the argument for open weights is that scrutiny makes models safer and better, the model should be easy to scrutinise. Mistral made theirs about as easy as it gets.
Sources
From the foundation