SimpleQA: measuring factuality with a benchmark models were meant to fail
OpenAI's new short-form factuality set was built by throwing out any question GPT-4 could answer, then graded with three labels instead of two. Notes on why the third label, not attempted, is the interesting part.
A dataset selected for failure
SimpleQA, released by OpenAI on November 7, is 4,326 short questions with single, verifiable answers, and the best model in the paper gets fewer than half of them right. That is by construction. When a trainer wrote a question, they were shown four completions from OpenAI models, mostly GPT-4 variants of different release dates, and the question was only kept if at least one of the four was wrong. Questions that every model answered correctly went in the bin.
The result is a set that sits at the edge of what the models know rather than in the comfortable middle. Older sets like TriviaQA and Natural Questions were useful once and are now saturated. SimpleQA was tuned to the current frontier the way an eye chart is tuned to the patient, and the authors say plainly that they expect it to stay relevant only for the next few generations.
The questions themselves are narrow on purpose. Trainers were told each one must have a single indisputable answer, must not change over time, and must be answerable as of December 31, 2023. So instead of asking where two people met, the question asks in which city. Instead of asking who someone's partner is in a long-running TV show, it names the season. Dates make up 32.8 percent of the answers and people another 24.1 percent, which tells you the flavour: obscure, checkable, and unlikely to be guessed from context.
How they know the answers are right
A factuality benchmark with wrong reference answers is worse than useless, so the verification step is worth describing. Every question was answered independently by a second trainer who could not see the first trainer's answer, and only questions where the two agreed were kept. The two trainers also had to cite sources from at least two different web domains between them.
After the set was frozen, a third trainer answered a random sample of 1,000 questions. The autograder scored them at 94.4 percent. The authors then read all 56 disagreements by hand. Fifteen were grader mistakes. Thirteen were the third trainer answering incompletely or contradicting their own cited source. The remaining 28, about 2.8 percent, were genuine problems with the questions, things like two reputable sites disagreeing on a date or a question with two correct answers. The paper's stated error rate is about 3 percent, and we appreciate that they show the arithmetic rather than just asserting the number.
Three grades instead of two
Every answer is graded by a prompted ChatGPT classifier as correct, incorrect, or not attempted. Correct means the response fully contains the reference answer without contradicting it. Incorrect means it contradicts the reference in any way, even if the contradiction is hedged. Not attempted means the reference is not given and nothing contradicts it, which covers both a plain refusal and an evasive answer that sends the user to a search engine.
This is the design decision that makes the benchmark worth writing about. A two-label grader treats a confident wrong answer and an honest abstention as the same failure. Once you split them, you can compute two numbers: overall correct, which is the share of all questions answered right, and correct given attempted, which is the share of attempted questions answered right. The authors combine them with a harmonic mean into an F-score, and acknowledge the loophole in that choice. If a model is below 50 percent, it is always better off guessing when it is at least half sure, so the F-score still mildly rewards attempting.
The paper offers an alternative that closes the loophole: score correct as one point, not attempted as zero, and incorrect as minus p, for a penalty p you choose based on how costly a wrong answer is in your setting. At p equal to 9, a model only scores positive if it gets 90 percent of what it attempts right. No model in the paper clears that bar.
What the table shows
GPT-4o answers 38.2 percent correctly, attempts 99 percent of questions, and is wrong on 60.8 percent. Claude 3.5 Sonnet answers 28.9 percent correctly but declines 35 percent of questions, so it is wrong on only 36.1 percent. Its accuracy on what it does attempt, 44.5 percent, is higher than GPT-4o's 38 percent. The two models end up with similar F-scores through opposite strategies, and which one you would rather deploy depends entirely on what a wrong answer costs you.
The small models make the point more bluntly. GPT-4o-mini attempts 99 percent of questions and is wrong on 90.5 percent of them. o1-mini declines 28.5 percent, gets a similar number right, and is wrong on 63.4 percent. Same knowledge, very different behaviour when the knowledge runs out. o1-preview is the best model in the table at 42.7 percent correct with 9.2 percent not attempted.
The authors note that Claude models abstain far more than the GPT-4o models, and that Claude still scores low, which they take as evidence the questions are hard in general rather than merely hard for the model they were filtered against. That is a fair sanity check, though it would be stronger with a model from a third lab.
Calibration as the real target
The second half of the paper uses the dataset to ask whether models know what they know. Two methods. First, ask the model for a confidence percentage alongside its answer and plot stated confidence against actual accuracy. Second, sample the same question 100 times at temperature one and plot how often the most common answer appears against how often it is right.
Both plots show a positive slope, so confidence carries some information. Both also sit well below the diagonal on the stated-confidence side, meaning every model overstates its confidence, and larger models are better calibrated than their smaller siblings. On the frequency method, o1-preview is close to the diagonal. A model that gives the same answer eight times out of ten is right about eight times out of ten. That is the property abstention depends on, and it is the one we would watch across model generations.
The limitation the authors put last is the one we would put first. This measures factuality on short queries with a single answer. Whether a model that abstains well here also abstains well in the middle of a 600-word essay containing thirty claims is an open question, and nothing in this paper touches it. What we would try next is running the same three-label grading over claims extracted from long answers, and seeing whether the not-attempted rate transfers or vanishes.
Sources
From the foundation