BrowseComp: when the human baseline is 29 percent
OpenAI's new browsing benchmark was built backwards, from answers to questions nobody can find again. The people who wrote it solved under a third of it.
Built from the answer outward
BrowseComp is 1,266 questions, each with a short answer that can be checked against a reference string. Jason Wei, Zhiqing Sun, Spencer Papay and seven coauthors at OpenAI released it this month as a benchmark for agents that browse the web. The design principle is stated in the paper's first pages. Answers should be hard to find and easy to verify.
The construction process is the interesting part. Trainers did not start with a question. They started with a seed, a person, an event or an artefact, then looked for several characteristics of it that each have a large search space, and wrote a question whose constraints only intersect at that seed. The paper's example is a request for an EMNLP paper whose first author went to Dartmouth and whose fourth author went to the University of Pennsylvania. Once you have the paper, checking takes a minute. Finding it by brute force means reading author pages for a whole conference.
Here is one from the benchmark itself. Identify the fictional character who occasionally breaks the fourth wall, has a backstory involving help from selfless ascetics, is known for his humour, and had a TV show that aired between the 1960s and 1980s with fewer than 50 episodes. The answer is Plastic Man. None of the individual clues is obscure. The intersection is.
Three filters and a two-hour clock
Every question had to pass three checks before it went in. GPT-4o with and without browsing, o1, and an early version of Deep Research all had to fail it. The trainer had to run five simple Google searches and confirm the answer was not on the first page of any of them. And another person had to fail to solve it within ten minutes, though the paper says that last rule was not enforced strictly.
Then they measured people. Trainers who had not written a given question attempted 1,255 of them with a two-hour limit. They solved 367, which is 29.2 percent, and gave up on 888. Of the questions that were solved, the trainer's answer matched the reference answer 317 times out of 367, or 86.4 percent. So even the solvable subset has a measurable disagreement rate, which is a useful number to keep in mind when reading model scores in the low single digits.
It matters who these humans are. They are the same population that wrote the questions, working in the same tooling, with the same idea of what a good search looks like. The 29 percent figure is a baseline for skilled searchers under a clock. It is not evidence that the questions are unsolvable, only that they are expensive.
From 0.6 percent to 51.5 percent
GPT-4o without tools scores 0.6 percent. GPT-4o with browsing scores 1.9 percent. o1, which has no browsing but reasons longer, gets 9.9 percent. Deep Research, the agent built to browse persistently, gets 51.5 percent. That ordering tells you the benchmark rewards two separate things. A browsing tool on its own barely moves the needle. Reasoning without a tool moves it more, presumably because o1 can recall and cross-reference more of what it has already seen. Combining persistent browsing with reasoning is what actually produces a score.
The 51.5 percent should also be read against the 29.2 percent human number with the caveats above. Deep Research had no two-hour limit and the humans did. The paper reports that Deep Research accuracy rises smoothly as test-time browsing effort increases, and that aggregating 64 samples per question with majority or confidence-weighted voting adds 15 to 25 percent over a single attempt, with best-of-N doing better still. Models are fairly good at recognising when they have the right answer on this task, even when they struggle to find it.
The calibration result cuts the other way. Models with browsing were worse calibrated than models without. GPT-4o with browsing had a calibration error of 82 percent and Deep Research 91 percent. The authors' reading is that access to web tools makes a model more confident in wrong answers. That is a real cost, since an agent that browses and then reports a confident wrong answer is harder to catch than one that says it does not know.
What the design principle buys and what it costs
Easy to verify, hard to find is a good principle for a benchmark that has to be graded automatically. Short reference answers mean no rubric, no judge model and no argument about partial credit. It also means the benchmark can be saturated without anyone getting confused about what saturation means. If an agent solves 90 percent of these, it can find entangled facts on the open web better than the people who hid them.
The authors are clear about what the principle excludes, and we want to repeat their list because it is easy to skip. BrowseComp does not sample from real user queries. It cannot measure long-form answers. It does not test whether an agent can resolve an ambiguous request. It measures single targeted retrievals rather than broad research tasks. And because the questions were constructed from one seed, the authors cannot guarantee that no other valid answer exists, which is a known failure mode of inverted construction.
The topic mix is also skewed toward things with well-indexed fan communities. TV and film is 16.2 percent of questions, science and technology 13.7 percent, art 10 percent, history 9.9 percent and sport 9.7 percent. That is where entangled but verifiable facts are cheap to write. It is not obviously where browsing agents will be used.
What we would do with it
We would treat BrowseComp as a persistence test rather than a research test. The 0.6 to 51.5 spread between a plain model and an agent is the clearest evidence we have so far that browsing agents are a different capability from browsing tools, and the test-time scaling curves make it a useful yardstick for how much compute an agent should spend before answering. The calibration numbers are the part we would keep watching. A benchmark that shows an agent getting more confident and wronger at the same time is doing its job.
The follow-up we want is a version of the human baseline with no clock and a different population. If professional researchers with a day per question also stall near 30 percent, the questions are as hard as the paper claims. If they reach 80, the human number is telling us about time pressure rather than difficulty, and the comparison to Deep Research needs to be redrawn.
Sources
From the foundation