445 benchmarks, few of them valid: the construct validity audit
Forty-two authors and 29 reviewers read 445 language model benchmarks and found that only 16 percent report any uncertainty, over a fifth never define what they measure, and nearly every paper had a weakness somewhere. Reading notes on what psychometrics would ask and the eight-item checklist the paper offers.
How the review was done
A team of 42 authors led by Andrew Bean, posted to arXiv on November 3 and appearing in the NeurIPS 2025 Datasets and Benchmarks track, went through 445 benchmark papers and scored each against the standards a psychometrician would apply to a test. The corpus started at 46,114 articles from ICML, ICLR, and NeurIPS between 2018 and 2024 and from ACL, NAACL, and EMNLP between 2020 and 2024. Keyword filtering cut that to 2,189, manual screening to 522, and 445 survived as benchmarks proper.
Twenty-nine expert reviewers did the coding. Mean inter-rater agreement was 0.524 on the Brennan-Prediger kappa, which is moderate. We mention that because a paper about measurement validity should be held to its own standard, and the authors report the number rather than hiding it.
Defining the thing being measured
Construct validity is the question of whether a test measures the concept it claims to measure. The first requirement is that the concept be defined. In this sample 78.2 percent of benchmarks define the phenomenon, which means 21.8 percent measure something they never say what it is. Among those with a definition, 52.2 percent use one that is widely agreed upon and 47.8 percent use one that is contested. 61.2 percent define the phenomenon as a composite of several sub-components and 36.5 percent as a single unified capability.
The composite case is where the trouble usually starts. If reasoning is defined as several sub-skills, then a single score averages across them, and a model can move the score by improving one sub-skill while the paper's claim is about the whole. Nothing in a headline number tells you which happened.
Where the items come from
The sampling numbers describe how benchmark items are chosen. 40.7 percent of benchmarks use constructed tasks and 28.5 percent use only constructed tasks. 33.6 percent draw all items from a single source. 43.3 percent handcraft new items, while 42.6 percent reuse items from existing benchmarks and 38.2 percent reuse data from earlier benchmarks or exams. 12.3 percent rely exclusively on convenience sampling and 27.0 percent use it in part, against 55.2 percent that use targeted sampling at least partially.
Reuse is the contamination pathway. Items lifted from an older benchmark or an exam have had years to appear in crawled training data. The paper's recommendation is to prepare for contamination, meaning to plan for it in the design rather than to check for it afterward, which is the reverse of current practice.
The 16 percent
The number we expect to be quoted from this paper is that only 16.0 percent of benchmarks use uncertainty estimates or statistical tests when comparing models. Every leaderboard gap smaller than the sampling noise of the test set is, on this evidence, being reported as a difference in 84 percent of cases without anyone checking. The related figures are that 53.4 percent present any evidence for construct validity, 35.2 percent compare against similar benchmarks, 32.4 percent include a human baseline, and 31.2 percent compare against a realistic setting.
The authors use GSM8K as a worked case, and they are fair about it. It is better than most. Their concerns are that the tasks may be confounded by reading comprehension, that the arithmetic may require a calculator to isolate the skill being claimed, and that there is no error analysis, so a wrong answer cannot be attributed to a cause. If that is the critique of a good benchmark, the bad ones are in worse shape, and the paper's summary is that nearly every paper had a weakness in at least one area.
The checklist
The eight recommendations are, in the paper's order: define the phenomenon, measure only the phenomenon, construct representative datasets, acknowledge the limits of reusing datasets, prepare for contamination, use statistical methods to compare models, conduct error analysis, and justify construct validity. None of these is novel to anyone who has built a test in psychology or education. What is new is applying them to a field that has mostly skipped them.
What we would want from the next version is a scored leaderboard of benchmarks themselves, using this rubric, so that when a lab reports a number the reader can see how much the number can bear. And we would like to see the 16 percent revisited in a year. If a paper this large with this many authors does not move that figure, nothing will.
Sources
From the foundation