BullshitBench: does the model push back on a broken premise?
Arena asked over 80 models 100 questions that sound expert and mean nothing, across software, finance, law, medicine and physics. Detection ranged from 2 percent to 91 percent, and turning on extended reasoning often made it worse. Notes on premise-checking as its own capability.
A question that cannot be answered correctly
Some questions have no right answer because the question is wrong. Arena's BullshitBench, written up by Peter Gostev on March 18, is built entirely out of those. Each item uses real terminology from one of five domains, software, finance, legal, medical, and physics, arranged in a structure that looks like something a practitioner would ask, and that on inspection is incoherent. The only good response is to say so.
The v2 set has 100 questions constructed with 13 distinct nonsense techniques, and the team ran it against more than 80 models from every major provider. Three judge models grade each response into one of three bins. Green means the model clearly pushed back on the premise. Amber means it partly recognised the problem. Red means it accepted the nonsense and answered it. The judges agree with each other about 80 percent of the time, which is decent for a three-way rubric and worth remembering when reading small differences.
The spread
The top of the table is Claude Sonnet 4.6 at 91 percent detection. The bottom is GPT-4o-mini at 2 percent. Between them, Qwen 3.5 scored 78, and GPT-5.4 and Gemini 3 Pro both scored 48. A range from 2 to 91 on the same 100 items is not the kind of spread you see on accuracy benchmarks, where frontier models cluster within a few points of each other. Whatever this measures, the labs are not all optimising for it.
The two models at 48 percent are the case we find most informative. They are strong on ordinary tasks and they accept a broken premise roughly half the time. So the ability to answer hard questions correctly and the ability to notice that a question is malformed are, at least at present, separable. A model can have a great deal of one and not much of the other.
Reasoning makes it worse
The result that should bother people is what happens when extended reasoning is switched on. For several model families, enabling the reasoning mode reduced detection. The author's hypothesis is that reasoning training optimises models to solve whatever problem they are handed, and a model that has been rewarded thousands of times for finding a path to an answer will find one here too, elaborately, for a question that has no answer.
This fits a pattern we have seen in our own work with chain-of-thought models. Given an inconsistent specification, the model rarely stops to say the specification is inconsistent. It picks the reading that makes the task solvable and proceeds, and the longer it reasons the more committed it becomes to that reading. Extra thinking tokens buy more confidence in a framing the model chose in the first few hundred.
Why this is a sycophancy eval in disguise
Most sycophancy evaluations work by stating an opinion and checking whether the model agrees. That measures a narrow thing, deference to an explicit stance. BullshitBench measures something broader, which is whether the model will defer to the implied authority of a question. Nobody in these prompts says they are an expert. The vocabulary says it for them, and a model that has learned to match the register of its user will answer in kind.
That is why the adversarial construction matters. If you ask a model a coherent question with an expert tone, agreement and correctness coincide and you learn nothing about which the model is tracking. You only separate them by building questions where going along with the user is the wrong move. The 13 nonsense techniques are, in effect, 13 ways of making agreement and correctness diverge, and the score is the fraction of the time the model chose correctness.
Here is a worked example of the sort of thing we mean, invented by us rather than taken from the benchmark, since the items themselves are held back. A user asks how to configure a database index so that a write transaction can be rolled back after it has been committed. Every term is real. The request is contradictory, since committed means it cannot be rolled back. A green answer says that and asks what the user is actually trying to protect against. A red answer produces configuration steps.
What we would want measured next
The obvious weakness is that the judges are themselves models, and a model that grades premise-checking may share the blind spots of the models it grades. An 80 percent agreement rate leaves plenty of room for correlated error, and we would like to see a human-graded subset before treating the ranking as settled.
The finding we most want followed up is the reasoning effect. If it holds across providers, then the training recipe that produces better math and code scores is actively degrading a different capability, and nobody would know from the leaderboards that reward solving. A benchmark where the right answer is to refuse the task is one of the few instruments that can see that trade-off. We would want every reasoning model release to report a number on it alongside the usual ones.
Sources
From the foundation