Chatbot Arena's paper: 240,000 votes and the statistics behind the leaderboard
The LMSYS team wrote up how the Arena leaderboard is estimated, from Bradley-Terry coefficients and sandwich standard errors to active sampling and a 600 cluster topic model of the prompts. A method deep dive on what the ranking measures and where its uncertainty hides.
What the data is
The Arena leaderboard has been cited by every major lab for months without a paper behind it. On March 7 Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Angelopoulos, Tianle Li and colleagues at LMSYS and Berkeley published one. The dataset it describes is more than 240,000 pairwise votes from around 90,000 users, collected between April 2023 and January 2024, across more than 50 models. Users type a prompt, two anonymous models answer, and the user picks the better one or calls a tie. Prompts came in over 100 languages, with English at 77 percent and Chinese at 5 percent.
That design choice is the whole method. Nobody wrote the questions, so there is no test set to leak into training data. Nobody assigned reviewers, so the judgments are as varied as the people who showed up. The paper's job is to argue that this uncontrolled pile of votes can be turned into a ranking you should believe, and to say how wide the error bars are.
Bradley-Terry instead of Elo
The leaderboard was launched on Elo scores, and the paper explains why the team moved away from them. Elo updates ratings sequentially, which is right for chess players whose strength changes over time and wrong for models whose weights are fixed. The Bradley-Terry model fits a single strength coefficient per model such that the probability of A beating B is a logistic function of the difference in coefficients. It uses all the votes at once, and it comes with standard statistical machinery.
For uncertainty, the authors compared bootstrap intervals with sandwich standard errors and settled on the sandwich version, which behaved better at their sample sizes. The published confidence intervals are what let you see that two adjacent models are often statistically tied. A one point gap on the leaderboard between models with overlapping intervals is not a result, and the paper is explicit that the ranking should be read with the intervals attached.
Which pairs to sample next
Votes are not free. Each one costs two model calls and a user's attention, and the interesting comparisons are between models that are close. The paper describes an active sampling rule that picks model pairs in proportion to how much a vote on that pair would shrink the confidence intervals. Against uniform random pairing, this cut the data needed to estimate the win matrix by 54 percent and the data needed for the scores themselves by 5 percent.
The gap between those two numbers is instructive. The win matrix has one entry per pair, so directing votes at uncertain pairs helps a lot. The Bradley-Terry scores pool information across every pair a model has played, so they were already reasonably efficient. If you only care about the top line ranking, sampling cleverness buys less than you would expect.
Are the prompts any good
The standard objection to crowdsourced evaluation is that users ask trivial questions and every model answers them fine. The authors ran topic modelling over the prompts and found 600 clusters spanning poetry, coding, mathematics and medical queries, with the largest cluster accounting for only 1 percent of the data. They then built a benchmark from a sample of prompts, called Arena Bench, and showed that it separates models, with a visible gap between proprietary and open models.
The other objection is that anonymous voters are not experts. The paper measured agreement between crowd votes and expert raters at 72 to 83 percent, and compared that with expert to expert agreement of 79 to 90 percent. The remaining 10 to 20 percent of disagreement is attributed largely to prompts with no single correct answer. Crowd voters are a little noisier than experts and a lot more numerous, and on this evidence the trade is a good one.
Where the uncertainty hides
The confidence intervals capture sampling noise. They do not capture the things we would worry about more. Preference is a property of the user population, and that population is whoever visits the site, which skews toward people who evaluate models for a living and ask the kind of questions they ask. Preference also rewards whatever voters respond to, including length and formatting, and the paper does not yet separate style from substance. A model that writes longer, tidier answers will climb, and the interval around its score will be small and honest about the wrong thing.
The paper addresses vote manipulation with a sequential test built on valid p-values and Fisher's combination method, reporting 90 percent true positive and 60 to 70 percent true negative rates on anomalous users. That is a start rather than a defence. As labs come to care about their Arena placement, the incentive to vote for your own model or against a rival grows, and 60 percent specificity means a lot of honest users will be flagged along the way. What we would want next is a per-topic leaderboard with its own intervals, a style adjustment that is reported alongside the raw score, and a public record of how many votes were excluded and why. The dataset is being released, so much of that can be done from outside.
Sources
From the foundation