What was released

On July 8 Hugging Face published SmolLM3, a 3 billion parameter decoder trained on 11.2 trillion tokens, in base and instruct variants. Elie Bakouch, Carlos Miguel Patiño, Anton Lozhkov, Edward Beeching and a couple of dozen coauthors wrote the announcement. The model is fine. It beats Llama-3.2-3B and Qwen2.5-3B on the base benchmarks and is competitive with the 4B Qwen3 and Gemma3. That is not why we are writing about it.

The reason is the list of artefacts underneath the model card. The nanotron training configs with the exact data weights for every stage. The training logs on wandb. The intermediate checkpoints. The SmolTalk2 dataset with its mid-training, SFT and preference subsets. The training scripts in nanotron, datatrove and lighteval. The post says the goal is to hand over the engineering blueprint, and gives the reason in one line: usually achieving these results would require months of reverse engineering.

Most open-weights releases give you the weights and a paper. Some give you a data description. Very few give you the configuration file that was actually loaded on the cluster, and almost none give you the loss curve. SmolLM3 gives all of that, and it is worth being concrete about what each piece lets someone else do.

The recipe, stage by stage

The architecture is a standard transformer with a few choices that are only interesting because they are documented. Grouped query attention with four groups to shrink the KV cache. Rotary position embeddings removed from every fourth layer, a variant called NoPE, to help long-context extrapolation. Intra-document masking so tokens do not attend across document boundaries in a packed sequence. Weight decay removed from the embedding layers for stability. Tied embeddings. Each of these is a small decision someone would otherwise have to guess.

Pretraining ran in three stages with published mixtures. Stage one, tokens zero to 8 trillion, was 85 percent web, 12 percent code, 3 percent math. Stage two, 8 to 10 trillion, shifted to 75, 15 and 10. Stage three, 10 to 11.1 trillion, went to 63 percent web, 24 percent code and 13 percent math. The global batch was 2.36 million tokens at a 4096 sequence length with a 2e-4 learning rate, on 384 H100s for 24 days. Those are the numbers you need to estimate what a comparable run would cost you.

Then two mid-training stages. A long-context extension of 100 billion tokens took the context from 4k to 32k to 64k, with YaRN used to reach 128k at inference. A reasoning stage of 35 billion tokens drew on OpenThoughts3-1.2M and the Llama-Nemotron datasets. Post-training was supervised fine-tuning on 1.8 billion tokens across 22 datasets for four epochs, one billion non-reasoning and 0.8 billion reasoning, followed by Anchored Preference Optimization using Tulu3 preferences for the non-reasoning mode and synthetic pairs from Qwen3-32B and Qwen3-0.6B for the reasoning mode.

The part you only learn from the logs

Here is the detail that justifies releasing logs and checkpoints rather than a summary. After preference optimisation the model had lost long-context ability. The fix was a linear merge, 0.9 of the APO checkpoint with 0.1 of the mid-training checkpoint, which recovered RULER performance at 64k. That is a repair, not a design, and it is the kind of thing that never makes it into a methods section. If you only had the final weights you would not know the merge happened. If you had the final weights and the paper you might know it happened but not what the alternative looked like. With the intermediate checkpoints you can reproduce the regression yourself and try a different fix.

The same applies to the reasoning mode. The instruct model supports a think and a no_think flag in the system prompt, and the no_think mode works by pre-filling an empty think block. In reasoning mode the model scores 36.7 percent on AIME 2025 against 9.3 without, and 30.0 on LiveCodeBench against 15.2. Those deltas come from the 35 billion token reasoning stage plus the SFT split. With the configs you can ask which of those two stages did the work, by training from the published checkpoint before each.

What reproducibility actually means here

Nobody outside a handful of labs is going to rerun 384 H100s for 24 days to check a loss curve. So the value of the release is not bit-for-bit reproduction. It is that every claim in the post is attached to an artefact you can inspect, and every stage can be resumed from a checkpoint rather than from scratch. A group with a modest cluster can take the stage-three checkpoint and test a different code fraction. A group with one node can take the mid-training checkpoint and try a different preference method. That is reuse, and it is worth more than reproduction.

It also changes what a comparison means. When two 3B models differ by two points on a benchmark, the usual explanation is data, and the usual response is a shrug, because the data mixtures are unpublished. With SmolLM3's mixtures published, the next 3B release can state its mixture relative to this one, and the two-point gap becomes an experiment rather than a mystery.

What we would like other labs to copy

Not the model. The habit. The minimum we would ask of any open-weights release is the config that was actually run, the loss curve, and one checkpoint per training stage. That costs almost nothing extra and it converts a release from a product into a piece of evidence. The multilingual and long-context claims in this post are only interesting because we can check them against the mixture and the extension schedule.

The experiment we would run first is the merge. Take the published APO checkpoint and mid-training checkpoint, sweep the merge weight from 0 to 1, and plot RULER at 64k against AIME. Hugging Face picked 0.1 and reported that it worked. The published artefacts make it possible to find out whether 0.1 was the right number or just the first one that did.

Sources

  1. Hugging Face blog: SmolLM3, smol, multilingual, long-context reasoner, July 8, 2025