Why this paper exists

EleutherAI's Language Model Evaluation Harness, lm-eval, has been the default way to score an open model for about three years, and it sits behind the Open LLM Leaderboard. Last week thirty of the people who built and maintain it, led by Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Leo Gao and Jonathan Tow, published a paper on what the job taught them. It is less a methods paper than a field manual, and its purpose is to write down folk knowledge that has so far lived in GitHub issues and Discord threads.

The problem they describe is simple to state. Two labs report a score on the same benchmark for the same model and get different numbers, and neither can tell why, because neither reported enough detail to find the difference. The paper's contribution is a catalogue of the places that difference hides.

The prompt is the benchmark

The sharpest example is Table 1. The authors ran five pretrained models on ARC-Challenge zero-shot under two prompt conventions. In the cloze style, the model is shown the question and its log-likelihood is scored for each answer's text. In the MMLU style, the choices are listed with letters and the model is scored on the letter. GPT-NeoX-20B scores 38.0 percent under cloze and 26.6 percent under MMLU style. Falcon-7B goes from 40.2 to 25.9. Mixtral-8x7B goes the other way, from 56.7 to 81.3.

So the choice of format can move a score by more than twenty points, and the direction of the move depends on the model. That is worse than a constant offset, because it changes rankings. A model that was trained on data resembling lettered multiple choice will look far better under one convention and a model that was not will look far worse, and a table that mixes the two conventions is comparing nothing.

The MMLU case study makes the same point at the level of whole implementations. The authors cite work by Clémentine Fourrier and colleagues comparing three independent MMLU implementations, HELM, lm-eval and the original code, which produced widely different scores and even changed the order of models. Part of the cause was that the Llama papers had quietly adopted the MMLU-style prompt for ARC as well, which nobody could tell from the reported numbers alone.

The small choices that are never written down

The paper lists a handful of benchmark quirks that each move results and that almost no paper reports. ARC is almost entirely four-choice, but one question has five options. MMLU should be micro-averaged over questions rather than macro-averaged over its 57 subjects, and the two differ by several percentage points. HumanEval has three problems that lack the example tests the other 161 include. None of these is a trick. Each is the kind of thing an implementer decides once, forgets, and never mentions.

Tokenization is the sneakiest of these. Most scoring code tokenizes the prompt and the candidate answer separately and concatenates the token lists, and the paper points out that widely used tokenizers give no guarantee this matches how the model would have tokenized the joined string. lm-eval works around it by moving trailing whitespace from the prompt into the target string. The general lesson the authors draw is to treat the model and its tokenizer as one system under evaluation and never assume the seam is harmless.

The same appendix walks through log-likelihood normalization. If answer choices differ in length, a raw log-likelihood favours the shortest one, so the harness also reports acc_norm, which divides by the byte length of each choice to remove the tokenizer from the comparison. Two papers reporting acc and acc_norm for the same task are reporting different things, and the labels often go missing in tables.

Point estimates that lie

The section we would make required reading is on variance. The authors reproduce a pair of plots from a post by Kamilė Lukošiūtė showing GPQA scores for several OpenAI and Anthropic models. As single-run numbers, the models sort into a clean order. With 95 percent confidence intervals from ten runs, the intervals overlap and the order dissolves. The same is true of HumanEval, which has only 164 problems, where a one point gain is inside the noise from sampling again at the same temperature.

Their point is that language models are stochastic objects and that most reported evaluations pretend otherwise. lm-eval reports bootstrapped confidence intervals alongside the score. The wider practice of reporting an isolated number with no interval, no seed and no run count is the reason so many leaderboard moves turn out to be nothing.

The rules we are taking from it

The paper's recommendations are unglamorous and we agree with all of them. Publish the evaluation code, and if you cannot, publish the exact prompts. Record the harness version and the commit hash. Say whether a multiple choice score is acc or acc_norm, and whether MMLU is micro or macro averaged. Report a confidence interval or at least the variance across runs. Never copy a number from another paper into your comparison table, because you do not know what setup produced it. And release the model outputs, so someone can rescore them later without rerunning the model.

What we would add is a habit rather than a rule. Before believing any gap between two models, rerun both under the same harness, same prompt style and same version, and look at the interval. A surprising share of the gaps that matter to a paper's story do not survive that check, and this paper finally gives us a citation for why.

Sources

  1. Biderman, Schoelkopf, Sutawika, Gao, Tow et al., Lessons from the Trenches on Reproducible Evaluation of Language Models (arXiv 2405.14782)