The table that carries the argument

Two OpenAI models on SimpleQA. The older one, o4-mini, abstains on 1 percent of questions, gets 24 percent right and gets 75 percent wrong. A newer thinking model abstains on 52 percent, gets 22 percent right and gets 26 percent wrong. On accuracy the older model wins by two points. On error rate it loses by 49. Any leaderboard that ranks on accuracy alone reports the first comparison and hides the second.

That is the whole argument of the paper by Adam Tauman Kalai, Ofir Nachum, Santosh Vempala and Edwin Zhang, and we think it is correct. A model that never says it does not know will beat a model that sometimes does, on a binary-graded test, as long as its guesses land above zero. Guessing is free under that scoring rule. The company post puts it plainly, that accuracy-only scoreboards dominate leaderboards and model cards, and that most benchmarks pluck out the accuracy metric and discard everything else.

Where the errors come from before anyone grades them

The paper splits the problem into two stages. In pretraining, it reduces generative error to a binary classification problem, is this output valid or not, and derives a bound on how well any model can do. The consequence is the singleton result: the hallucination rate is at least the fraction of facts that appear exactly once in the training data. If a person's birthday shows up in one document and nowhere else, no amount of scale makes that fact reliably retrievable, because there is no statistical signal separating it from a plausible alternative.

The second stage is post-training, and that is where the paper puts the blame for the behaviour we actually see. A model that has learned it cannot be sure could express that, and reinforcement from evaluations teaches it not to. The paper lists widely used benchmarks and notes that their scoring gives no credit for an admission of uncertainty. A confident wrong answer and an abstention score the same, which is zero, so the model that guesses collects the expected value of guessing and the honest model collects nothing.

This is a familiar structure. Standardized tests worked out decades ago that unpenalized wrong answers make guessing optimal, and some of them added negative marking to fix it. The paper is asking our field to do the same arithmetic.

The proposed fix is one line in the prompt

The specific proposal is to state a confidence target inside the evaluation prompt. The form given is to answer only if you are more than t confident, since mistakes are penalized t divided by one minus t points, while correct answers earn one point and an admission of uncertainty earns zero. At t equal to 0.75, a wrong answer costs three points, so answering is worth it only when the model believes it is at least 75 percent likely to be right.

Two things about this design are worth noticing. It is explicit, so the model is not left guessing what threshold the grader has in mind, and different thresholds can be published side by side to show how a model behaves under different costs of being wrong. And it requires no new dataset. An existing benchmark can be rescored under this rule with the same questions and the same answers, as long as the scoring code records abstentions instead of counting them as failures.

The company post frames the same idea as a scoring change: penalize confident errors more than uncertainty, and give partial credit for appropriate expressions of uncertainty. Both framings put the burden on the evaluation rather than on a separate hallucination benchmark, and the post is explicit that adding another hallucination eval alongside the accuracy scoreboards will not fix the incentive.

What would count as adoption

The obvious question is whether any leaderboard that people actually watch will change. We have no evidence yet of a major benchmark adopting confidence targets, and the paper does not claim one. So the honest thing to say in September 2025 is that the proposal exists and the incentive it describes has not moved.

What adoption would look like is concrete enough to check later. The scoring code would need an abstention detector that does not simply pattern match on the phrase we do not know, since a hedged wrong answer is still a wrong answer. A leaderboard would need to publish the threshold it scored at, because a model tuned for t equal to 0.5 will look bad at 0.9. And model cards would need to report the error rate next to the accuracy, which costs nothing and would already have exposed the o4-mini comparison above without any new machinery.

The part we are least sure about is whether the singleton bound survives retrieval. The result concerns facts seen once in pretraining data, and a model that can look something up is not sampling from its priors. If the bound is a statement about parametric recall rather than about deployed systems, then the scoring fix matters mainly for closed-book evaluations, and the field should say which regime each benchmark is testing. We would rather see that spelled out than see another eval added to the pile.

Sources

  1. OpenAI: Why language models hallucinate
  2. Kalai, Nachum, Vempala and Zhang, Why Language Models Hallucinate (arXiv 2509.04664)