Adding error bars to evals
Evan Miller's paper for Anthropic argues that most reported benchmark differences come without a confidence interval and many would not survive one. A walkthrough of the five recommendations, from treating questions as a sample to running a power analysis before you build the eval.
Questions are a sample
Anthropic published a short paper on November 19 by Evan Miller called Adding Error Bars to Evals. Its starting move is one that experimental scientists take for granted and eval reports almost never do. The 1,000 questions in a benchmark are a draw from a much larger population of questions that could have been written, and the number we care about is the model's skill on that population rather than its luck on this draw. Once you say that, everything else in the paper follows from ordinary statistics.
The first recommendation is to report a standard error of the mean from the central limit theorem and to put a 95 percent interval of plus or minus 1.96 standard errors around every score. Bootstrapping works too but is not needed. What this buys you is the ability to say whether a two-point lead between two models is a lead at all, and the paper's point is that a surprising fraction of the leads in published tables are not.
Clusters make it worse
The second recommendation is the one most people will not have thought of. Many benchmarks have questions that are not independent. A reading comprehension eval asks several questions about one passage. A multilingual eval asks the same question in several languages. If the model finds one passage hard, it gets all of that passage's questions wrong together, and the naive standard error, which assumes independence, understates the uncertainty.
The numbers are large. On DROP the clustered standard error is 3.05 times the naive one, 1.34 points against 0.44. On MGSM it is 1.88 times. On RACE-H, where the clusters are smaller relative to the set, it is 1.10 times. A difference between two models on DROP that looks like three standard errors under the naive calculation is one standard error under the correct one, and that is the difference between a finding and noise.
Where the variance comes from
The third recommendation splits the variance of a score into two parts. Some questions are hard and some are easy, and that spread is the variance of the conditional means. On top of that, a model sampled at nonzero temperature gives different answers to the same question, and that is the conditional variance. Only the second kind can be reduced by sampling more. If you ask each question K times and average, the conditional part falls by a factor of K, and in the paper's uniform-difficulty example going from one sample to two removes a third of the total variance.
Two corollaries are worth stating. For multiple choice without chain of thought, reading the next-token probabilities instead of sampling removes the conditional variance entirely. And lowering the temperature does not reduce noise. It moves variance from the conditional part into the conditional means, which no amount of resampling touches. People who set temperature to zero to get stable evals are getting a stable number, which is a different thing from an accurate one.
Compare in pairs
The fourth recommendation is that when two models are run on the same questions, you should analyse the paired differences rather than two independent means. Question difficulty is shared, so the difference cancels it. The paper reports that frontier models' per-question scores correlate at somewhere between 0.3 and 0.7, and at a correlation of 0.5 the paired analysis cuts the variance of the difference by a third. In the worked example, paired standard errors reverse the apparent ordering of two fictional models on one benchmark, which is exactly the sort of thing that happens in a leaderboard row.
Decide the size before you build it
The last recommendation is a power analysis. Before writing an eval, decide the smallest difference you would want to detect, and compute how many questions that requires. The formula is the standard one with the paired variance in the numerator and the effect size squared in the denominator. The paper's example is that detecting a three-point difference at 80 percent power and a 5 percent false positive rate takes about 969 questions, which is why the recommendation is that new evals have at least a thousand items.
That number is a problem for the field, because many of the benchmarks that matter most right now have a few hundred questions, and a few dozen questions in the hard subsets people quote. What we would like to see is every leaderboard rendering a confidence interval, and every new benchmark paper stating the minimum detectable effect. If a benchmark cannot separate models that differ by five points, saying so in the abstract would save a lot of arguments.
Sources
From the foundation