Why LLaMA's MMLU score depended on who ran it
The Open LLM Leaderboard gave LLaMA 65B an MMLU of 0.488. Meta reported 0.636 and HELM measured 0.637. Hugging Face traced the gap to three implementations of the same benchmark that differ in prompt format and how the answer is read out.
A 15 point gap on the same model and the same questions
MMLU is 57 subjects of four-way multiple choice. Nothing about it is ambiguous. Yet LLaMA 65B scored 0.636 under the original implementation from the authors of the benchmark, 0.637 under Stanford CRFM's HELM, and 0.488 under the version of the Eleuther AI evaluation code the Open LLM Leaderboard was running. Same weights, same questions, same five-shot setting, a gap of about 15 points. Hugging Face published the investigation on June 23.
The interesting part is that none of the three implementations is wrong. Each made a defensible choice about two things that MMLU never specified: what exactly to put in the prompt, and how to turn the model's output into one of four letters.
Three ways to read an answer
The original implementation compares the model's probability for each of the four letters, A, B, C and D, and takes the highest. It does not include a Question prefix and does not name the subject. HELM generates text and matches it against the full answer strings, includes a Question prefix and adds formatting spaces. The Eleuther AI implementation computes the log-likelihood of each complete answer sequence rather than the letter, omits the topic, and prepends a Choices keyword.
Scoring the letter and scoring the full answer text are different measurements. The first asks whether the model can produce the label. The second asks whether the model assigns high probability to the answer sentence, which mixes in how long and how fluent that sentence is. Longer answers are penalized by an unnormalized log-likelihood, so an implementation that scores full answers will systematically favour models whose calibration on longer strings happens to be better. Generation with string matching adds a third confound, which is whether the model will emit the answer in the expected surface form at all.
The prompt differences push the same way. A topic line tells the model this is a question about anatomy, which is a real hint. A Question prefix cues the format. A Choices keyword changes what the model expects next. None of these are large edits to a prompt, and together they were enough to move a 65B model by 15 points.
The rankings moved too
If every implementation shifted every model by the same amount, this would be a footnote. It does not. The post shows Falcon 7B scoring above LLaMA 7B under the Eleuther implementation and below it under both HELM and the original. A reader comparing two models on a leaderboard was getting an answer that depended on which implementation the leaderboard used, and that dependency was not disclosed anywhere in the number.
This is what made the leaderboard controversy of the last few weeks worth having. People assumed a discrepancy meant somebody had cheated or misreported. The actual explanation was that three research groups implemented the same paper independently and made different reasonable calls, and that the community had been treating benchmark names as if they identified an experiment. MMLU names a dataset. It does not name a measurement.
What to report alongside a score
The lesson we would take is narrow and practical. When you report MMLU, report the implementation, the commit, the shot count, the prompt template and the answer extraction method. The Eleuther AI commit in question was e47e01b, and naming that commit is the difference between a reproducible claim and a rumour. A score without that metadata is not comparable to any other score, including your own from three months ago.
The same applies when reading someone else's. If a model card cites a number from a paper and a number from a leaderboard in the same table, those numbers are probably from different implementations and the table is not a comparison. This is easy to check and almost nobody does it.
Hugging Face resolved this one by aligning the leaderboard code with the original implementation, and the Eleuther AI maintainers made the change while the post was being written. That fixes MMLU. It does not fix the general problem, which is that most benchmarks specify a dataset and leave the measurement procedure to the reader. We would like to see benchmark releases ship reference evaluation code with a pinned version and a conformance test, so that a second implementation can be checked against the first before anyone publishes a leaderboard built on it. Ten questions with published expected scores would be enough to catch every difference described above.
Sources
From the foundation