What Arena shipped

Arena's post today describes a change to how a new model gets its first number. Until now a model entered the arena and waited days for enough anonymous pairwise votes to accumulate before it received a rating. Under AutoEval, a reward model trained on Arena's own preference data casts votes automatically, and the model gets a provisional score in under an hour. The score is labelled AutoEval on the board and is replaced as human votes come in.

The reward model is pointwise, mapping a single prompt and response to a scalar, and it was trained on what Arena describes as millions of pairwise comparisons from live votes, with training data running through the end of April. There are versions for text, vision, image generation and coding. For images the training set exceeds three million preference pairs, and Arena reports the model leads the MMRB2 benchmark by more than nine points.

The numbers behind the claim

The headline figure is a rank correlation above 0.98 between AutoEval rankings and the live human rankings. Arena also reports that the text reward model predicts held-out human preferences 8 to 10 percent more accurately than frontier language model judges, naming Gemini 3 Flash and Pro and GPT-5, and that in head-to-head model comparisons it picks the right winner more than 90 percent of the time when the true gap exceeds 10 points and every time beyond 15 points.

Those are good numbers for a surrogate, and we want to be precise about what they measure. A rank correlation of 0.98 across a board of many models is dominated by the models that are far apart. It says almost nothing about the ordering of the five models clustered at the top, which is the only ordering anyone reads the board for. The head-to-head figure is the more honest one, and it says the surrogate is reliable when the gap is large and unreliable, by its own account, when the gap is small. Small gaps are where the new models land.

A model of the voters

Here is the question we keep returning to. Arena's value was that its scores came from people. A reward model trained on those people's past votes is a model of what they used to like. Arena's own list of open challenges says as much. Human preferences shift as expectations rise, model capabilities keep advancing and require sensitivity to smaller differences, preference labels carry genuine ambiguity and ties need explicit modelling, and multi-turn feedback is hard to attribute. Every one of those is a way the surrogate can drift from the population it stands in for.

The drift matters more because of who reads the score. A lab shipping a model checks the board within the hour, and the number it sees is the reward model's opinion. If the human votes disagree a week later, the press release has already gone out. That is a mild version of the problem. The strong version is that a lab can now train against AutoEval directly, because a pointwise reward model trained on Arena votes is exactly the artifact you would use as a reward signal in post-training. Arena has not released the weights, so the attack is indirect, but a black-box scalar you can query at scale is enough to optimise against.

Goodhart's law in this setting has a specific shape. The reward model learned features that predicted votes through April. Some of those features are quality. Some are style, length, formatting, and confidence, the same confounders that made length-controlled AlpacaEval necessary two years ago. A model that climbs the AutoEval score by improving quality also climbs the human score. A model that climbs it by matching the confounders climbs the provisional score and then drops when the humans arrive. The gap between the provisional and final scores for each model is therefore the most interesting number Arena could publish, and the post does not mention it.

What would make us comfortable

The design decision we agree with is the label. Marking scores as AutoEval and replacing them with human votes keeps the two signals separate, and as long as the board never averages them the human leaderboard is still a human leaderboard. The provisional score is a forecast of it, and forecasts are fine if you score them.

So score them. Publish, for every model that has passed through the provisional phase, the AutoEval rating, the eventual human rating, and the difference, along with the date of the reward model's training cutoff. Retrain the reward model on a schedule and version it on the board, so that a score from July 2026 can be distinguished from one produced by the model that will replace it. And report the head-to-head accuracy at gaps of 2 and 5 points, since those are the gaps that decide who is first.

The wider point is that every evaluation that started as human judgment is drifting toward a model of human judgment, because humans are slow and models are cheap. Arena is doing this more carefully than most. The question of whether the result is still a human preference leaderboard depends on how visible the substitution is, and right now the answer is visible enough to trust and not yet measured enough to check.

Sources

  1. Arena, AutoEval scores