Six ways evaluation is harder than it looks: Anthropic's challenges post
Anthropic's policy post walks from multiple-choice benchmarks to third-party audits and reports what went wrong at each step, including a five-point swing on MMLU from formatting alone and a bias score of zero that meant the model was refusing to answer.
The claim about MMLU
Anthropic published a post on October 4 about the difficulty of evaluating its own models, and the detail that has been travelling is the one about formatting. Their words: simple formatting changes to the evaluation, such as changing the options from (A) to (1) or changing the parentheses from (A) to [A], or adding an extra space between the option and the answer, can lead to a roughly 5 percent change in accuracy on MMLU. That is a benchmark with 57 subjects whose reported scores are compared across labs to a tenth of a point.
They list three further problems with MMLU on top of that. The questions leak into training data, so a model may have seen them. Labs implement the eval differently, few-shot or not, chain of thought or not, so cross-lab comparisons are not comparing the same thing. And some questions are mislabelled or unanswerable. None of these is new to people who run evals, but we have not seen a lab put a number on the formatting effect before, and five points is larger than most of the differences that get reported as progress.
A bias score of zero
The second example is the one we would make people read. BBQ is a benchmark for social bias in question answering across nine social dimensions. Anthropic found no working open-source implementation and reports that it took one of their best full-time engineers one uninterrupted week to implement and test it. The bias score runs from minus one, anti-stereotypical, through zero, no bias, to one, stereotypical.
Some of their models scored zero. The post says this made them feel optimistic about progress on bias, until they looked and found the models were not answering the questions at all. In their phrasing, the results were technically unbiased and completely useless. A metric had been built to detect one failure and was silent on a different failure that produced the same number. If you take one lesson from the post it should be that any score which can be achieved by refusing to engage needs a second score next to it that measures engagement.
Third-party suites: too big or too slow
The post then covers the two large community evaluation frameworks. BIG-bench has 204 evaluations from more than 450 contributors. Anthropic reports that installing it took significant engineering effort, that it did not scale well, that they found bugs in implementations including the BBQ-lite variant, and that working out which of the 204 tasks were representative would have been a research project of its own. They ran it once and did not continue.
HELM has the opposite profile. It is curated top-down by experts, so there is little engineering to do, but iteration is slow and evaluating a new model can take months. There was also a format mismatch. Claude models were trained on a Human and Assistant dialogue format, HELM prompts differently, and the post says the resulting numbers gave a misleading impression of the model's performance. The two frameworks fail differently and neither fails in a way that is fixed by more compute.
Humans, red teams and model-written evals
Human A/B tests, where crowdworkers pick the more helpful or harmless of two responses in open-ended dialogue, come with their own problems. They are expensive and slow, they depend on the creativity and motivation of whoever is judging, and they expose the tension that a model can look harmless by being unhelpful. The post says plainly that there is no clear numerical threshold for how harmless is harmless enough.
Red teaming for national security risks is where the difficulties compound. It requires expert and sensitive knowledge. The post calls red teaming more art than science and says there is no standard way to make results comparable. Security clearances restrict what domain experts and developers can share, and a model that produces controlled information may create legal liability for the lab that elicited it. Model-generated evaluations are faster, minutes instead of days or months, but the models inherit the biases and fabrications of the models that wrote them, and humans still have to verify the results, which brings back the human problems.
The audit that needed the auditee
The last level is a third-party audit, in this case with the Alignment Research Center on dangerous capabilities. Anthropic expected it to be straightforward and reports that it required significant science and engineering support on their side. The structural problem is that the auditor has to keep distance to stay credible, while the lab holds most of the practical knowledge of how to get behaviour out of the model, so the audit is either less independent or less effective than you would like. That is an honest account of a tension that will not go away.
Their policy asks follow from all this. Fund the science of repeatable evaluations through grants to computer science departments. Fund implementations of existing good evals so they are not rebuilt from scratch at every lab. Fund work on how fragile existing evals are. Give NIST more money and consider a public safety leaderboard on the model of NIST's Face Recognition Vendor Test. And create legal safe harbours and disclosure protocols so that national security evaluations can be done and their results shared.
What we would add
The post is a list of problems from a lab that also publishes leaderboard numbers, and we read it as an argument against trusting any single number, including theirs. The two examples that stick are the ones with a figure attached, five points from brackets and a zero that meant silence. Both are cheap to check on any eval you run. Vary the prompt format and report the spread. Count the non-answers separately from the wrong answers.
What we want from the field next is for those two checks to become reporting conventions, in the same way error bars are a convention elsewhere. A benchmark score with a format sensitivity and a refusal rate beside it would have prevented both of the embarrassments in this post, and it would make the next lab's number comparable to this one's, which right now it mostly is not.
Sources
From the foundation