The setup

On 3 May, Lianmin Zheng, Ying Sheng, Wei-Lin Chiang, Hao Zhang, Joseph Gonzalez and Ion Stoica at LMSYS launched Chatbot Arena. The interface is two chat windows side by side. You type a prompt, two anonymous models answer, you vote for the better one or declare a tie, and only then are the model names revealed. Votes cast after the names are shown do not count.

Between 24 April and 1 May they collected 4,700 valid anonymous votes across nine open models: Vicuna-13B, Koala-13B, OpenAssistant's Pythia-12B, Alpaca-13B, ChatGLM-6B, FastChat-T5-3B, Dolly-v2-12B, LLaMA-13B and StableLM-Tuned-Alpha-7B. Prompts were mostly English. The votes are converted into Elo ratings, the system chess uses, where each pairwise outcome shifts the ratings of the two players by an amount that depends on how surprising the result was, using a base-10 logistic curve for the expected win probability.

Why pairwise and why Elo

The motivation in the post is that static benchmarks were failing for chat models in two ways. Multiple-choice academic tests do not measure whether a model is helpful in an open conversation, and the alternative of using GPT-4 as a judge, which this same group introduced with Vicuna in March, has known biases. Humans comparing two answers to their own prompt is closer to the thing anyone actually wants to know.

Pairwise comparison is the right primitive for that because absolute scores from untrained raters are noisy and drift. Two answers to the same prompt are easy to rank. Elo is the natural aggregator because it handles incomplete comparison graphs, gives a running estimate without waiting for a full tournament, and produces a number people already know how to read. The first leaderboard has Vicuna-13B at 1169, Koala-13B at 1082, OpenAssistant Pythia at 1065, and Alpaca-13B at 1008, with the rest below 1000 down to StableLM at 858.

The design choices they had to make on the fly

Two decisions in the post are worth noting because they will shape how the numbers should be read. First, the initial matchmaking paired models according to a guess at their strength, so that close matches would be more informative. They later switched to uniform sampling. The consequence is that some pairs fought many more battles than others, and models added late have fewer games. The authors flag this explicitly.

Second, they chose not to release the conversation logs, citing privacy and toxicity concerns. That is a reasonable call for a week-old project, but it means nobody outside LMSYS can yet audit what kinds of prompts drive the ratings. A leaderboard where the test set is invisible is only as trustworthy as its operators, and the authors say they plan to address this.

They are also candid that Elo here is a relative measure. It tells you Vicuna beat Koala more often than not. It does not tell you that either is good at anything in particular, and the post promises task-specific rankings as a later addition.

What the caveats predict

Reading the limitations section as a set of forecasts, we would bet on the following. The non-uniform sampling problem will get fixed quickly, because it is a matter of code. The visibility problem will recur every time a new model climbs the board, because people will want to know what the voters were asking. And the missing closed models will not stay missing: the post lists ChatGPT, GPT-4 and Claude as planned additions, and once they are in, the leaderboard becomes the de facto public scoreboard for the whole field rather than for open models.

The caveat we would add that the authors did not is about who votes. 4,700 votes in a week from whoever found the link is a self-selected population, and its prompts will reflect what that population cares about. Elo will faithfully rank models by how well they please that crowd. Whether that is the ranking anyone else wants depends on who shows up next.

What we would like to see

The obvious next experiment is to release the prompts, or at least a categorised sample, so that the rating can be decomposed by task. The second is to report confidence intervals rather than point Elo values, because 4,700 votes spread over 36 possible pairings is not many per pair. If those two things happen, this could become the first evaluation of chat models that is both continuous and hard to overfit. If they do not, it will be a popularity contest with a good user interface.

Sources

  1. Zheng et al., Chatbot Arena: Benchmarking LLMs in the Wild with Elo Ratings (LMSYS blog, 3 May 2023)