What was disclosed and when

FrontierMath launched on 8 November 2024 as a set of hundreds of original research level mathematics problems, written and checked by mathematicians, with automated verification and the explicit promise that the problems were unpublished to minimise contamination. At launch the best models solved under 2 percent. On 20 December OpenAI used the benchmark in the o3 announcement, and on the same day Epoch AI disclosed that OpenAI had supported the benchmark's creation.

Epoch's follow up statement this month fills in what "supported" means. OpenAI commissioned Epoch to produce 300 problems that form the core of the benchmark, and OpenAI owns all 300. OpenAI has the statements and solutions to 250 of them. For the remaining 50, a holdout set, OpenAI receives the problem statements but not the solutions. Epoch says the agreement did not prevent it from telling contributors that an AI company was sponsoring the work, and that its communication with them should have been more systematic and transparent.

The contributors' side came out through a LessWrong post by an Epoch contractor and through Carina Hong, a Stanford mathematics PhD student, who told TechCrunch that six mathematicians she spoke to were unaware OpenAI would have exclusive access, and that most said they would not have contributed had they known. Tamay Besiroglu of Epoch said the organisation made a mistake and should have negotiated harder for the ability to be transparent with contributors. Elliot Glazer, Epoch's lead mathematician, said of OpenAI's reported scores that Epoch cannot vouch for them until its independent evaluation is complete.

Why this is a measurement problem first

The ethical failure is obvious and Epoch has conceded it. What we want to separate out is the measurement failure, because it survives even if you believe every assurance about how the data was used. A benchmark is a claim that a number means something. FrontierMath's meaning rested on two properties, that the problems were unpublished and that no lab had seen them. Once one lab has seen 250 of 300 problems with solutions, the second property is gone for that lab, and the number OpenAI reports has a different meaning from the number anyone else reports.

It does not matter, for this purpose, whether OpenAI trained on the problems. What matters is that nobody outside OpenAI can distinguish the case where it did from the case where it did not. A 25 percent score on a set you have read is not comparable to a 2 percent score on a set you have not, and a benchmark on which the comparison is not possible has stopped being a benchmark for the lab in question. Epoch's statement that it cannot vouch for the number until it runs its own evaluation is the right position, and it is also an admission that the December headline was, at the time, unverifiable.

The 50 problem holdout is the fix, and it is a partial one. Fifty problems gives you a wide confidence interval, and the holdout only exists because the sponsor agreed to it. The general lesson is that access is a property of the benchmark, not of the lab's intentions, and it needs to be published with the score.

Humanity's Last Exam and the other way to fund a benchmark

Humanity's Last Exam, posted on 24 January by the Center for AI Safety and Scale AI with over a thousand co-authors, is the contrast case in the same month. It is 2,500 questions across dozens of subjects, 24 percent multiple choice, the rest exact match, about 14 percent requiring an image. The questions were collected from nearly a thousand contributors across more than 500 institutions, with a $500,000 prize pool and $5,000 for each of the top 50 questions.

The filtering is the interesting part. Over 70,000 candidate questions were run against frontier models, roughly 13,000 that the models could not answer went forward to expert review, and 2,500 survived. That process has a known weakness, which the authors acknowledge by construction. A question that stumps the models of January 2025 is selected for exactly that, so the benchmark is adversarial against a specific generation and the scores of that generation are biased low.

With that caveat, the launch numbers are still striking. GPT-4o scored 2.7 percent, Claude 3.5 Sonnet 4.1, Gemini 1.5 Pro 4.6, o1 8.0. On the text only subset DeepSeek-R1 got 8.5 and o3 mini high got 13.4. Calibration error was between 73 and 89 percent across the board, meaning the models were confidently wrong on almost everything they got wrong. The paper's own abstract says popular benchmarks like MMLU are now above 90 percent and no longer informative, and HLE is offered as a closed ended exam that is not.

What independence actually requires

Both benchmarks are expert written, contamination conscious and expensive. The difference is who paid and what they got for it. HLE's money came in as prizes to contributors and the questions went public. FrontierMath's money came from a lab that then owned the questions and could read most of them. Both approaches have a cost. Public questions will be trained on, and by next year HLE will need its own holdout. Private questions funded by a lab are, for that lab, no longer private.

We do not think the answer is that labs should never fund evaluations. They are the ones with the money and the interest. We think the answer is that the funding and access terms are part of the result and have to be reported alongside it, in the same table, every time. A score from a lab with solution access should carry a marker that says so, and a score from the holdout should be reported separately with its own interval. Epoch is now doing something like this. It would have been better if the benchmark had launched that way.

What we would like to see tried is a standing convention, adopted by benchmark maintainers rather than labs, of publishing an access table with each release. Who funded it, who has seen which subset, and what the holdout is. That is a small amount of paperwork, and this month shows what it costs to skip it.

Sources

  1. Clarifying the creation and use of the FrontierMath benchmark (Epoch AI, Jan 2025)
  2. AI benchmarking organization criticized for waiting to disclose funding from OpenAI (TechCrunch, Jan 2025)
  3. FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI (Epoch AI, Nov 2024)
  4. Humanity's Last Exam (Phan et al., Jan 2025)