Length-controlled AlpacaEval: fixing the judge that loved long answers
AlpacaEval win rates could be moved by 40 points with a single prompt asking for more detail. Dubois and colleagues regressed length out of the judge and the correlation with human votes went up. Notes on the fix and what it says about every model-judged leaderboard.
How badly the judge could be gamed
AlpacaEval asks a GPT-4 Turbo judge to compare a model's answer against a baseline answer on 805 fixed instructions and reports the fraction of wins. It is cheap, it runs in minutes, and until this week it had the highest correlation with Chatbot Arena of any automatic benchmark, at 0.94 Spearman. It also had a problem everyone knew about and nobody had fixed. The judge liked long answers.
The new paper from Yann Dubois, Balazs Galambosi, Percy Liang and Tatsunori Hashimoto measures how bad that was. They took several models and ran them three times, with a standard prompt, with a prompt asking for as much detail as possible, and with a prompt asking to be as concise as possible. GPT-4 1106 preview, which is the baseline and therefore should score 50 by construction, scored 22.9 when told to be concise and 64.3 when told to be verbose. Claude 2.1 went from 9.2 to 24.4 on the same manipulation. A single sentence in the system prompt could move a model past several competitors.
The fix is a regression
The authors treat length as a mediator in a causal graph. The model identity affects the judge's preference directly, which is the effect we want, and it also affects output length, which in turn affects the preference, which is the effect we do not want. The counterfactual question is what the win rate would be if the model's outputs were the same length as the baseline's.
To answer it they fit a logistic regression on the existing judge decisions with three terms. One is a per-model coefficient. One is a tanh of the standardised length difference between the model output and the baseline output, with a per-model slope. One is a per-instruction difficulty term shared across models. The length-controlled win rate is then the model's predicted win rate with the length term set to zero. Because the formula is antisymmetric in the two models and the length term vanishes when lengths match, the corrected score keeps the properties of a win rate. A model against itself still scores 50 and swapping model and baseline still gives 100 minus the score.
One detail matters for anyone who wants to copy the method. The instruction difficulty term is estimated once across all models and then frozen, and each new model's coefficients are fit separately reusing it. That means adding a model to the leaderboard does not change anyone else's number, which is the property that makes a public leaderboard usable.
What the correction did
Gameability fell. GPT-4 1106 preview now ranges from 41.9 to 51.6 across the three verbosity prompts instead of 22.9 to 64.3, and the normalised standard deviation across prompts dropped from 25 percent to 10 percent. Spearman correlation with Chatbot Arena rose from 0.94 to 0.98, computed on the 38 models present in both, which the authors say makes it the automatic metric with the highest agreement with human votes that they know of.
The leaderboard itself reshuffled in a telling way. Proprietary models, which tend to write shorter answers, gained. GPT-4 0613 gained 14.4 points of win rate and 20 places. The biggest losses went to open models that had been through preference optimisation, including several PairRM and DPO variants with average output lengths above 2,500 characters. The authors put this politely, saying the results are consistent with open models having exploited the length bias. We would put it more directly. The length bias had been selecting for models that were tuned to the judge rather than to people.
They also checked the correction against a cheap adversarial attack, truncating every output to a few characters except on the instructions where the model already does well. A naive length correction rewards that attack, lifting GPT-4's score from 3.7 to 25.9. Weak regularisation on the length slope brings it back to 12.2 without affecting non-adversarial models. It is not a complete defence, and they do not claim it is.
The warning underneath
The paper is about AlpacaEval, but the mechanism is general. Any leaderboard scored by a model has confounders that the model finds easier to detect than the quality it is supposed to measure, and as soon as people optimise against the leaderboard those confounders get optimised. Length was the one everyone could see. The authors list others from prior work, including lists in the output and position in the comparison. Each of those can be regressed out the same way, but only once someone names it.
What we take from the paper is that a model-judged score is a measurement with a known bias term, and a leaderboard should publish the correction it applies and the residual it cannot. The 0.98 correlation is good news. The fact that it took a year of open models climbing the old leaderboard before anyone measured the bias is the part we would want the field to remember. The next spurious correlate is presumably already being optimised somewhere, and the way to find it is to run the same verbosity style manipulation on whatever feature you suspect and watch whether the score moves.
Sources
From the foundation