Style control: what happens to the Arena when you subtract markdown and length
LMSYS added four style covariates to its Bradley-Terry model and reran the leaderboard. Some models fell a dozen places. Notes on the method, the coefficients, and how much of a human preference vote is presentation.
The suspicion
Everyone who has used Chatbot Arena for long has had the thought: the model on the left won because it wrote more, used headers, and put the answer in a bulleted list. Human raters like a response that looks organised. The question LMSYS set out to answer this week is how much of the leaderboard that preference explains, and the answer is enough to move models by up to twelve places.
The post went up on August 28 and introduces a style-controlled version of the leaderboard alongside the standard one. Nothing about the votes changed. What changed is the model that turns votes into rankings.
How the correction works
The Arena leaderboard is fit with a Bradley-Terry model, which is a logistic regression where each model has a strength parameter and the probability that A beats B is a function of the difference in strengths. Style control adds covariates to that regression. For each pairwise comparison, LMSYS computes the difference between the two responses in four features: token length, number of markdown headers, number of bold elements, and number of lists. Those differences enter the regression as additional terms with their own coefficients.
The effect is that any part of A's win rate that is explained by A being longer or more formatted gets attributed to the style coefficient instead of to A's strength parameter. What is left in the strength parameter is the part of the preference that the four style features cannot account for. The authors describe this as attributing increases in strength to the confounder, as opposed to the model, which is a clean way to say it.
This is standard causal-inference practice for observational data, and LMSYS is careful to say that is all it is. There could be confounders they did not measure. And the one they worry about most is that length can correlate with genuine quality, for instance when a longer answer contains reasoning steps that a shorter one skipped. The correction cannot tell that kind of length from padding.
What the coefficients say
When length and markdown are controlled together, the length coefficient is 0.249. Lists come in at 0.031, headers at 0.024 and bold at 0.019. So length dominates, and the formatting features are each an order of magnitude smaller. All four are positive, meaning more of each is associated with winning, holding model identity fixed.
The magnitude is what surprised us. A coefficient of 0.249 on normalised length difference is a substantial share of the strength gap between models that sit a few places apart on the board. It means a model can climb several ranks by being trained to answer at greater length, with no change in whether its answers are right.
Who moved
On the overall leaderboard, GPT-4o-mini dropped from rank 6 to rank 11 under style control. Grok-2-mini fell from 6 to 18. Claude 3.5 Sonnet rose from 6 to 4 and Claude 3 Opus from 16 to 10. On the hard prompts category, Claude 3.5 Sonnet moved from 2 to a tie for first, and Llama 3.1 405B moved from 4 to 3.
The pattern is consistent with what people had observed informally. The models that fell are the ones known for long, heavily formatted answers. The models that rose are the ones that answer more tersely and use less markdown. Whether the corrected ranking is closer to the truth depends on what you think the leaderboard should measure, and that is the real argument this post starts.
Which leaderboard is right
There are two defensible positions. One says the Arena measures what users prefer, users prefer formatted answers, and a correction that removes that preference removes real signal. If a product ships a model, the formatting is part of the product. The other says the leaderboard is used as a proxy for capability, and capability is what a lab can claim when it advertises a rank, so a rank that can be bought with verbosity is misleading.
We hold the second position for one reason. The style features are cheap to optimise. Any lab can add length and headers with a fine-tuning pass that costs nothing, and once one lab does it every lab has to. A leaderboard that rewards that converges on everyone producing the same long formatted answer, and the votes stop carrying information about anything else. Subtracting the style is a way of keeping the metric from being gamed to death.
What we would like to see next is the same regression with more features, particularly a measure of whether the answer contains explicit reasoning steps, so that the length coefficient can be split into reasoning length and padding length. If the padding coefficient is still 0.2 after that split, the case for the controlled board is settled. If it drops to near zero, then length was quality all along and the sceptics were right.
Sources
From the foundation