Search-time contamination: the agent that found the answer key on Hugging Face
Scale AI researchers logged search-enabled agents retrieving the benchmark they were being tested on, complete with labels. Blocking Hugging Face cut accuracy on the affected questions by about 15 percent. Notes on a contamination class that no amount of pretraining decontamination can fix.
What the logs showed
A paper from Scale AI posted on August 12 describes something we had assumed was happening but had not seen documented. Ziwen Han, Meher Mankikar, Julian Michael and Zifan Wang ran search-enabled agents on three benchmarks, Humanity's Last Exam, SimpleQA and GPQA, and read the retrieval logs. In about 3 percent of questions the agent pulled up a Hugging Face dataset page containing the very question it was answering, with the ground-truth label sitting next to it.
The agents did not hide this. The reasoning chains say, in effect, that a matching question and answer pair was found on Hugging Face. In one example the authors show in their first figure, Sonar Deep Research compares its own answer to the ground truth from an ungated HLE upload, notices the two disagree, and goes with the upload. That is a model that got the question wrong and was marked correct because the answer key was one search away.
Which agents and how often
The work concentrates on Perplexity's Sonar family: Sonar Pro, Sonar Reasoning Pro and Sonar Deep Research. The authors also tried agents from Anthropic, Google and OpenAI and found those almost never surfaced a Hugging Face link, which they attribute to narrower retrieval or filtering on the provider side rather than to any virtue of the underlying models. So the numbers below describe one retrieval stack, and a different stack could be better or worse.
The contamination rates by benchmark were roughly 3.3 to 3.4 percent on HLE, about 1 to 1.2 percent on SimpleQA, and 1.9 to 4.15 percent on GPQA depending on the agent. Small percentages, but concentrated on exactly the questions where a lookup helps most. On HLE with Sonar Pro the accuracy gap between contaminated and clean questions was more than ten points, and when the authors blocked Hugging Face from the search results accuracy on the contaminated subset fell by about 15 percent. That is the direct benefit the agent was getting from reading the key.
Why this is a new class of problem
Everything we normally do about contamination happens before training. Labs run n-gram overlap against the pretraining corpus, hold back canary strings, keep private test splits, and argue about whether a benchmark has been seen. Search-time contamination sidesteps all of it. The weights can be spotless. The leak happens at inference, through a tool call, against a live copy of the web that includes whatever someone uploaded last week.
It also gets worse over time rather than better. A benchmark released in January is clean in January. By March there are three mirrors of it on Hugging Face, a blog post with worked solutions, and a paper appendix with the answers. The authors' date-cutoff ablation found that pages published after a benchmark's release contribute a measurable share of the accuracy gain, and that a large amount of the useful evidence lives outside Hugging Face, in papers, blogs and mirrors. Blocking one domain plugs the most obvious hole and leaves the rest.
What the authors recommend
The recommendations are practical rather than clever. Give evaluators multiple search filters so they can block source sites and report which ones they blocked. Run an internal audit on retrieval logs using keyword filtering and substring matching against the benchmark text. Document mitigations in the results table, not a footnote. Prefer information-seeking benchmarks like BrowseComp and Mind2Web, where using the web is the task, over knowledge tests where the web is a shortcut. And release the complete evaluation logs so other people can find contamination you missed. Scale released theirs.
That last point is the one we would push hardest. A search agent's score without its retrieval trace is a number with no provenance. If we cannot see what it read, we cannot tell whether it reasoned or copied, and the paper shows that even the agent's own summary of what it did is sometimes the only evidence that it copied.
What we would do with this
For anyone publishing agent scores on HLE, GPQA or SimpleQA, the minimum now is a contamination check on the logs and a reported blocklist. For benchmark authors, we think the honest position is that gating a dataset on Hugging Face does very little once a single ungated copy exists, and that the semi-private test split model that ARC Prize uses is going to spread to knowledge benchmarks too.
The open question we would like someone to answer is how much of the gain from search on these benchmarks survives when every post-release page is excluded, not just the ones on one domain. The authors' ablation suggests the answer is a good deal less than the headline numbers imply, and that would change how much we trust any agent score with a browsing tool attached.
Sources
From the foundation