A Saturday release

Llama 4 arrived on April 5, a Saturday, two days before the LlamaCon date most people had expected. Nathan Lambert called the timing utterly bizarre and read it as a sign of a release pulled forward under pressure. Two models shipped. Scout is a mixture of experts with 16 experts, about 109 billion total parameters and 17 billion active. Maverick has 128 experts, about 400 billion total and the same 17 billion active. Both are natively multimodal through early fusion of text and image tokens, both were pretrained on up to 40 trillion tokens across 200 languages, and both are distributed under a custom community licence.

A third model, Behemoth, was announced as the teacher that Maverick was co-distilled from, still in training, and has not been released. Twelve days on, the mood in the open weights community is closer to disappointment than to the excitement that met Llama 3, and we want to separate what the models actually are from the way they were presented.

What the architecture changed

The interesting engineering is in how the models handle position. Meta calls it iRoPE. Every fourth layer drops rotary position embeddings entirely and attends over the full sequence with causal masking. The remaining layers use RoPE with chunked attention over 8,192-token windows, which keeps their memory cost bounded. The no-position layers get a temperature scaling on attention to counter the way attention probabilities flatten over very long inputs. Scout also normalises query and key states in the RoPE layers. The idea is that the chunked layers do local work cheaply and the sparse set of global layers carry long-range information without positional encodings that would have to be extrapolated.

Scout is full mixture of experts in every layer, while Maverick alternates expert layers with dense ones. Both share the 17 billion active figure, which is the number that determines per-token compute, and Scout's total size is what lets Meta say it fits on a single server-grade GPU with 4-bit or 8-bit quantisation applied on the fly. The training recipe used a hyperparameter transfer method Meta calls MetaP, described as inspired by muP, and Maverick was trained against Behemoth's logits with dynamic weighting. Hugging Face's benchmark table has Maverick at 80.5 on MMLU Pro and 69.8 on GPQA Diamond, with Scout at 74.3 and 57.2.

What 10 million tokens means

Scout's instruction-tuned model is listed with a 10 million token context, Maverick's with 1 million. Both were pretrained at 256 thousand. The gap between the pretraining length and the advertised length is filled by the iRoPE design and by whatever long-context post-training Meta did, and the evidence offered for it is needle in a haystack. Lambert's line is that needle tests are a necessary condition and not a sufficient one, and that the release did not report RULER or any of the newer long-context suites that measure aggregation and reasoning across a document rather than retrieval of a planted string.

We want to be precise about the objection, because it applies beyond this release. A context window is a number of positions the model will accept without error. Whether it can use them is a separate empirical question, and passing a needle test at 10 million tokens tells you that the global layers can route one fact from far away. It does not tell you that the model can summarise, compare or follow an argument across that span. Until someone publishes those numbers for Scout, the 10 million figure describes the input buffer and nothing else.

The Arena score and the community

The part that turned reception cold was the leaderboard. Meta's launch material put Maverick at 1417 on LMArena, second overall. The footnote disclosed that the entry was an experimental chat version tuned for conversation, and the weights that were actually released are a different model. Lambert called the results fake in the sense that matters, since the benchmarked model cannot be downloaded, and described it as a slight to the community that had built the Llama ecosystem. Whatever the intent, the effect was that the one number in the launch that people could compare against Gemini and GPT-4o turned out not to describe the artefact.

His broader argument is that Meta designed for the wrong audience. The people who made Llama the default open model were LocalLLaMA hobbyists, academics and mid-sized businesses who wanted a range of dense sizes they could run and fine-tune. A 109 billion parameter MoE that needs a server GPU and a 400 billion one that needs several is a different offer, and the small dense models that would have served the old audience did not ship. Lambert's verdict is that Llama is no longer the open standard, which was true the week Qwen and DeepSeek passed it and is now hard to dispute.

What we would test

None of this makes the models bad. Maverick's benchmark numbers are competitive for the active parameter count, and iRoPE is a real design that deserves independent study. The tests we want are concrete. Run Scout on RULER and on a long-document question set at 128 thousand, 1 million and 10 million tokens and publish the curve, so the context claim becomes a plot instead of a footnote. Then ablate the global layers, since if long-range performance collapses without them the architecture has taught us something about where long context lives.

The other question is about Meta. Behemoth is described as still in flight, and Lambert reports independent evaluation putting it behind Gemini 2.5 Pro. If it ships, the co-distillation story holds together. If it does not, then Llama 4 was two student models without their teacher, and the open weights community will spend the year on other people's checkpoints while waiting to learn what Meta does next.

Sources

  1. Hugging Face: Welcome Llama 4 Maverick and Scout
  2. Nathan Lambert, Interconnects: Llama 4: Did Meta just push the panic button?