What happened with Maverick

When Meta launched Llama 4 at the start of April, the Maverick model it cited on LM Arena was a version called Llama-4-Maverick-03-26-Experimental, which Meta described as chat optimised. That version did well. The weights Meta actually released, Llama-4-Maverick-17B-128E-Instruct, were evaluated by LM Arena separately after the discrepancy came out, and as of 11 April the released model sat in 32nd place, below GPT-4o, Claude 3.5 Sonnet and Gemini 1.5 Pro, several of which were months old.

LM Arena apologised, said Meta's interpretation of its policy did not match what it expected from model providers, and changed its rules. Meta's statement was that it experiments with all kinds of custom variants, that the experimental version was one of them, and that developers are free to customise the open source release. Both statements are accurate. Neither addresses the problem, which is that a leaderboard position was cited for a model nobody could download.

The ranking that people saw, the one that shaped the launch coverage, was for a model tuned to the arena's human raters. The ranking that describes the thing you can run is 32nd. If the point of a leaderboard is to tell a user what to expect, the first ranking was noise and the second was the signal, and the noise came first.

The paper that shows it was not a one off

On 29 April a group led by Sara Hooker at Cohere, with co-authors from Princeton, Stanford, Waterloo, Washington and elsewhere, posted The Leaderboard Illusion. It is a study of how Chatbot Arena's operating practices interact with its rankings, and the Maverick case is one data point in it. The authors identified 27 private variants that Meta tested on the arena before the Llama 4 release. The published leaderboard shows one of them.

That is the core mechanism. If a provider can test many variants privately and publish only the best, the published score is a maximum over samples rather than a single draw, and it is biased upward by an amount that depends on how many variants were tried. A provider who submits one model gets one draw. A provider who submits 27 gets the best of 27. The leaderboard does not record which is which.

The second mechanism is data. The paper estimates that Google and OpenAI received 19.2 percent and 20.4 percent respectively of all arena data, while 83 open weight models combined received 29.7 percent. Access to arena conversations lets a provider tune to the arena's distribution, and the authors estimate that limited additional arena data can produce relative performance gains of up to 112 percent on arena style evaluation. Their conclusion is that these dynamics produce overfitting to arena specific behaviour rather than general model quality.

What a human preference leaderboard is actually measuring

It is worth being precise about what LM Arena measures even when nobody games it. A rater sees two anonymous responses to their own prompt and picks one. The aggregate is a preference over responses on the distribution of prompts that arena users happen to type, judged by the arena users who happen to vote. That was always a proxy for quality, and a useful one, because it was hard to fake and it correlated with things people cared about.

The proxy breaks when the target knows about it. A model tuned to be preferred by arena raters is being tuned toward length, formatting, tone and confidence in the style those raters reward, and the paper's Maverick numbers are a demonstration that this tuning moves the ranking a long way without moving the released model at all. Once the benchmark is a launch metric, the incentive to do exactly this is strong, and the arena's own data access policies make it easier for the largest providers than for everyone else.

None of this makes the arena useless. It makes it a measurement with known biases, which is what every measurement is. The problem is that its position in the industry, as the number cited in launch posts, assumed it was unbiased, and the last month has shown that assumption was wrong in a way that favours the labs with the most variants to test and the most data to tune on.

What we would change

The paper's recommendations are the sensible ones and we will not restate all of them. The two that we think matter most are that every variant tested should have its score published, whether or not the provider releases it, and that data sampling across providers should be equalised and disclosed. Both are policy changes on the arena's side, and the arena has already changed its rules once this month, so they are feasible.

The change on the reader's side is to treat a leaderboard position as a claim about a specific artefact. If the artefact you can download is not the one that was ranked, the rank tells you nothing about what you are downloading. Meta's 32nd place is the honest number for Maverick, and we would rather have that with a clear label than a top five with an asterisk nobody reads.

What we would like someone to try is a simple audit. Take the released weights of every model in the arena's top twenty, run them through the arena's own pipeline under a fresh anonymous name, and publish the gap between the fresh ranking and the listed one. If the gap is small the leaderboard survives. If it is large, we learn how much of the top of the board is variant selection rather than model quality, which is the number this paper's methods can estimate but a direct experiment could measure.

Sources

  1. Meta's vanilla Maverick AI model ranks below rivals on a popular chat benchmark (TechCrunch, Apr 2025)
  2. The Leaderboard Illusion (Singh et al., Apr 2025)