Reproducing 2,200 ICML papers in nineteen days
A Hugging Face community challenge pointed coding agents at a third of ICML 2026. Half the papers had a claim verified, a quarter had one falsified, and 242 got contradictory verdicts from different teams. What agent-scale reproduction settles and what it cannot.
The scale of the thing
Between July 15 and August 2, 1,221 people attempted to reproduce 2,226 of the 6,352 papers accepted at ICML 2026. That is 34 percent of the conference in 19 days. They published 6,816 reproduction logbooks, launched 2,962 cloud jobs, and released 274 full agent traces. Each participant got 20 dollars of Hugging Face compute credit, which tells you the individual budgets were tiny and the scale came from the tooling.
The tooling was coding agents. Participants used Claude Code, Codex, Cursor, and Pi to read a paper, extract its core claims, write code, run experiments, and write up what happened. Every logbook was then judged automatically by GLM-5.2 into one of four verdicts, verified, falsified, toy, or inconclusive. So there are two layers of model in the loop, one doing the work and one grading it, and both matter for how to read the totals.
What the verdicts say
Of the 2,226 papers, 1,103, or 51 percent, had at least one claim verified. Within that group 266 were judged fully reproduced and 632 partially. Across all papers 3,978 individual claims were verified. On the other side, 496 papers, or 23 percent, had at least one claim falsified or contested, and 49 had every checked claim falsified. A further 502 papers produced only toy-scale evidence, and 280 were inconclusive.
Those percentages overlap, because a paper with several claims can land in more than one bin. The number that stops the totals from being a clean scorecard is 242. That is how many papers received contradictory verdicts from different teams, one saying a claim held and another saying it did not. Roughly one in nine of the attempted papers, then, has a reproduction status that depends on who tried.
The falsifications that look real
The write-up gives three examples where the disagreement with the paper is specific enough to check. A paging algorithm paper claimed a competitive ratio of H_k plus a constant. The reproduction measured something closer to H_k plus a term growing like log k, which is a different asymptotic claim. A theorem about attention and Frank-Wolfe optimisation was tested numerically and counterexamples appeared at steps 224, 3,800, and 6,416. A self-distillation paper analysed the reverse KL divergence in its theory while the released code optimised the forward KL.
The third case is the one we would show a student. The paper is not wrong in the sense of a fabricated result. The theory and the code simply describe different objectives, and nobody noticed because nobody had lined them up. An agent that reads the proof and the training loop in the same session is well placed to notice that, and it takes no GPU at all.
What contradictory verdicts tell us
Take the 242 papers with split verdicts. There are at least three ways to get one. The agent picked a different claim to test, so the two teams did not check the same thing. The agent made an implementation choice the paper left unspecified, and the choice mattered. Or one team's toy-scale run happened to fall on the right side of noise. The automated judge cannot distinguish these from its seat, and neither can we from the summary.
That is the limit of agent-scale reproduction in its current form. It is very good at producing a first pass over thousands of papers and surfacing the ones where something concrete does not line up. It is not good at adjudicating. The challenge's own write-up says the most reliable results came from workflows where a human was steering, re-pointing the agent and questioning an assumption. The 6,816 logbooks are evidence that the agents did a lot. The 242 splits are evidence that the agents alone do not know what they found.
There is also the grader to think about. GLM-5.2 issued every verdict. If it systematically reads a partial match as verified, the 51 percent is inflated. If it reads any deviation as falsified, the 23 percent is inflated. A human-graded sample of a few hundred logbooks would tell us which, and we have not seen one yet.
What we think this changes
Before this summer, reproducing a third of a major conference was a multi-year project that nobody funded. Now it is a nineteen-day community event on 20-dollar credits. Even discounting the verdicts heavily, that changes what authors should expect. Somebody will run your code against your claims within weeks of acceptance, and the mismatch between your appendix and your training script will be found by a model that reads both.
What we would like the next round to do differently is to make the disagreements the product. Publish the 242 split papers as a list, assign two humans to each, and report how many splits were about claim selection, how many about unspecified details, and how many about noise. That is the dataset that would let us calibrate the automated verdicts, and it is the one thing an agent cannot generate for itself.
Sources
From the foundation