MT-Bench and the 80 percent agreement number
The paper that formalised LLM-as-a-judge reports GPT-4 agreeing with human experts 85 percent of the time, above the 81 percent humans manage with each other. The same paper measures position bias, a verbosity attack and self-preference. A close look at what the agreement figure licenses and what it quietly excludes.
What the paper built
Lianmin Zheng and twelve co-authors from the LMSYS group have written down the method everyone has been using informally since Vicuna in March, and then tested it. The paper contributes two datasets. MT-Bench is 80 handwritten two-turn questions across eight categories: writing, roleplay, reasoning, math, coding, extraction, STEM knowledge and humanities knowledge. Chatbot Arena is the crowdsourced platform where users chat with two anonymous models and vote, which collected around 30K votes in its first month. Both are released, along with about 3K expert votes on MT-Bench.
The judge comes in three flavours. Pairwise comparison shows the judge two answers and asks which is better or whether they tie. Single-answer grading asks for a score out of ten for one answer. Reference-guided grading gives the judge a reference solution, which the authors use mainly for math. The default judge throughout is GPT-4, and the paper's central question is whether its verdicts line up with what human raters say.
Where the 85 percent comes from
Table 5 is the source of the headline. Six models answered all 80 questions, expert human labellers each judged at least 20 random multi-turn questions, and the authors computed agreement as the probability that two randomly chosen judges of each type agree on a randomly chosen pairwise comparison. Under the setup they call S2, which keeps only non-tie votes, GPT-4 pairwise judging agrees with human experts 85 percent of the time on the first turn. Human experts agree with each other 81 percent of the time on the same subset. That is the sense in which GPT-4 matches the human-human rate.
The setup matters. S2 excludes every case where either judge declared a tie or where GPT-4 gave inconsistent verdicts when the answers were swapped. Under S1, which counts ties and treats inconsistent verdicts as ties, GPT-4 agreement with humans drops to 66 percent, and human-human agreement to 63 percent. Random agreement is 50 percent under S2 and 33 percent under S1. So the fair summary is that on the clear-cut comparisons, a GPT-4 judge is about as reliable as a second human. On the hard, close comparisons the paper reports a much lower number and the abstract does not mention it.
Position bias, measured
The bias section is the part we expect to be cited longest. To isolate position bias the authors generated two near-identical answers per first-turn question by calling GPT-3.5 twice and asked each judge to compare them in both orders. GPT-4 gave the same verdict in both orders 65 percent of the time and favoured the first answer 30 percent of the time. GPT-3.5 was consistent 46.2 percent of the time. Claude-v1 was consistent 23.8 percent of the time and picked the first answer 75 percent of the time, and when the assistants were renamed it turned out Claude also had a preference for the name Assistant A.
The proposed fix is cheap and we would treat it as mandatory. Run every comparison twice with the order swapped and only award a win when the same answer wins both times, otherwise call it a tie. Few-shot examples in the judge prompt raised GPT-4's consistency from 65 to 77.5 percent, though the authors note that consistency is not the same as correctness.
Verbosity and self-preference
The verbosity test is a small adversarial attack. Take 23 MT-Bench answers that contain a numbered list, ask GPT-4 to rephrase the list without adding information, and prepend the rephrased list to the original so the answer is twice as long and says nothing new. Claude-v1 and GPT-3.5 preferred the padded answer 91.3 percent of the time. GPT-4 fell for it 8.7 percent of the time. All three judges correctly tie two identical answers, so the judges can handle the format and still get the judgement wrong.
Self-enhancement bias is harder to isolate because you cannot easily get a model to grade an answer without knowing its own style. The authors settle for a statistical look. Relative to human win rates, GPT-4 as judge gave itself a win rate about 10 percent higher and Claude-v1 gave itself about 25 percent higher. GPT-3.5 did not favour itself. The math and reasoning results are the most sobering: on ten math questions GPT-4 judged an incorrect answer correct in 14 of 20 comparisons with the default prompt, 6 of 20 with chain of thought, and 3 of 20 when given a reference answer.
What the figure does and does not license
Here is how we would use this paper. If you are comparing chat models on open-ended tasks where the answers differ clearly in quality, a GPT-4 judge with order swapping is a defensible substitute for a panel of humans and roughly as consistent as one. If your comparison is between two strong models whose answers are close, the 66 percent S1 figure is the relevant one, and you should expect a third of your verdicts to be noise. If the task has a right answer, do not use a judge without a reference solution.
The two things the paper cannot tell you are what happens when the judged models are stronger than the judge, and what happens when models start being trained against the judge. Both are coming. GPT-4 favouring itself by 10 points is a small effect today because only one model is being judged by itself. Once every lab tunes on GPT-4 preference, the bias becomes the optimisation target. We would want a replication of Table 2 and Table 3 every time a new judge model ships, and we would want the S1 numbers reported alongside the S2 ones in anyone's leaderboard.
Sources
From the foundation