The numbers

The ICLR 2026 program chairs published their retrospective today, and the scale is the first thing to absorb. The conference received 19,525 valid, format-compliant submissions. It desk rejected 779 for procedural or content violations and saw 5,042 withdrawn, leaving 13,763 papers that went to a final decision. Those papers drew 76,139 reviews from 18,054 reviewers. In the end 5,355 were accepted and 8,408 rejected, an acceptance rate of 27.4 percent.

Read those figures as a systems problem. Eighteen thousand reviewers is a small city, recruited, assigned, and chased within a few months, and the number of reviews is larger than the entire submission count of most conferences a few years ago. Any process that runs at this size is going to be decided by tooling, enforcement, and failure modes more than by the taste of any individual area chair.

What the LLM policy actually required

ICLR 2026 required that any use of LLMs be disclosed, and it stated that individuals were responsible for their contributions regardless of what tools they used. Violations were treated as code of ethics matters rather than as style complaints. That framing matters, because it moves undisclosed model use out of the grey zone of writing assistance and into the same category as plagiarism or collusion.

The enforcement side is what we had not seen a conference do at this scale before. Two LLM content detectors were run across all of the more than 75,000 reviews, and flagged reviews were sent to area chairs to assess for quality. Submissions were also checked with automated reference verification against bibliographic databases, and papers with hallucinated references went through three stages of human review, first the area chair, then the program chairs, then an appeal channel. The retrospective does not report how many reviews or papers were flagged at each stage, and we would like to see those counts.

There is a tension here that the chairs do not fully resolve. Detectors of model-written text are unreliable on any single document, which is why the flagged reviews went to a human rather than being discarded. But sending a flag to an area chair who already has too many papers converts an automated signal into more unpaid work. The design is defensible. Whether it changed outcomes is unknown from what has been published.

The exploit

The event that defined this cycle was disclosed on December 3. An OpenReview API exploit made it possible to recover the supposedly anonymous names of authors, reviewers, and area chairs. The bug may have been known and used as early as November 11. Someone scraped author, reviewer, and other details for over ten thousand submissions, 45 percent of the conference, and distributed the data online. OpenReview found and fixed the bug on November 27, and ICLR reverted reviews on November 28 and 29.

The consequences the chairs describe go beyond privacy. With reviewer identities in the open, authors could collude with reviewers and third parties could pressure them. The December post states that reviewers were harassed, intimidated, and offered bribes. The retrospective confirms that reports of collusion, threats, and harassment were investigated, that the community members responsible were banned, and that their submissions were desk rejected.

How the process was rebuilt mid-cycle

The response was blunt because it had to be. Reviewer discussion was frozen immediately so that nothing further could be influenced. Every paper was reassigned to a new area chair, on the reasoning that an attacker with the leaked mapping would no longer know whom to lean on. Review text and scores were reverted to their state before the bug became known, which means that any rebuttal-period movement, legitimate or not, was thrown away. The meta-review period was extended to January 6, 2026.

Think about the cost of that decision. Authors who had earned a score increase through a good rebuttal lost it. Reviewers who had updated in good faith had their updates discarded. The chairs accepted a large amount of collateral damage to remove the possibility that a leaked identity had shaped a score. We think that was correct, and we also think it only works once. A community will tolerate one reset. It will not tolerate a review process where resets become a routine mitigation.

Security and logistics come first now

Put the two halves together. Nearly twenty thousand papers, eighteen thousand reviewers, detectors sweeping every review, reference checkers sweeping every bibliography, and in the middle of it a platform vulnerability that exposed almost half the conference. The intellectual work of peer review, a careful person reading a paper and deciding whether the claims hold, sat at the end of a pipeline whose integrity depended on an API, a bug bounty that did not exist, and a chair's willingness to reset everyone's scores.

The retrospective is honest about this in the way it allocates its attention, which is mostly to process, tooling, and enforcement. That is the right emphasis. The open questions we have are ones the community can answer with data the chairs already hold. How many reviews did the detectors flag, and how many of those did area chairs judge to be low quality? How many papers had references that did not resolve? Did the score reset change any decisions relative to the pre-incident scores? Publishing those numbers would turn one conference's emergency into a baseline the next one can plan against.

Sources

  1. ICLR blog, A retrospective on the ICLR 2026 review process (March 31, 2026)
  2. ICLR blog, ICLR 2026 response to security incident (December 3, 2025)