The numbers

The main track received 21,575 submissions this year and accepted 5,290, an acceptance rate of 24.52 percent that the chairs describe as consistent with prior years. For scale, the 2020 conference had 9,467 submissions, so the load has more than doubled in five years while the acceptance rate has been held almost fixed. That is a policy choice, and it means the number of accepted papers has more than doubled too.

Behind the decisions were 20,518 reviewers, 1,663 area chairs and 199 senior area chairs. The chairs say plainly that recruiting at this scale meant a larger share of first-time NeurIPS area chairs, and that papers might be reviewed and handled by less experienced reviewers, ACs and SACs. The Datasets and Benchmarks track, run separately, took 1,995 submissions on top of that, up from 1,820 in 2024, with its own committee of 41 SACs, 281 ACs and 2,680 reviewers. The conference was split across two venues this year, San Diego and Mexico City, and added a position paper track for the first time.

What they tried

The main structural change was what the chairs call responsible reviewing. Authors who had reviewing obligations and did not meet them had their own reviews withheld. Papers whose coauthors showed gross negligence in reviewing were desk rejected, and the chairs report eleven such cases. Eleven is small against 21,575 but the point was deterrence, and the chairs seem to think it worked well enough to keep.

The other change was to bidding. Concerns about fake accounts and collusion rings, where groups of authors arrange to review each other's papers, led the chairs to reduce reliance on reviewer bids and depend more on OpenReview's algorithmic matching. That trades one failure mode for another. Bids let a reviewer pick papers they understand. Matching picks papers the system thinks they understand. At 20,000 reviewers the second is probably the lesser evil, but it is not free, and it feeds directly into the experience problem above.

The rebuttal problem

The observation in the reflections we keep coming back to is about fatigue. Some authors engaged heavily in rebuttals, and the chairs report that some reviewers responded by raising their scores to end the discussion rather than by engaging with the substance. That is the review process failing in the quietest possible way. The score goes up, the paper is accepted, and nobody wrote down that the reviewer gave in.

The chairs' answer was to lean on consensus between ACs and SACs to counteract review noise, and to acknowledge that imperfect outcomes were inevitable given the resources. We read that as an admission that the reviewer score is no longer the primary signal. It is one input to a committee judgement, and the judgement is being made by people who are themselves stretched across more papers than they can read closely.

What the datasets track shows

The Datasets and Benchmarks track is a useful control because it tried something different. Submissions were required to host datasets on a recognised platform and to provide structured metadata, and reviewers received automated metadata reports and compliance checklists. In the post-conference survey of 851 authors and 155 reviewers, 77 percent of reviewers said datasets were easy to access, 69 percent found the automated reports useful and 70 percent used the checklists. Eleven percent of reviewers said they did not inspect the dataset directly, which is a number we would like to see go to zero.

The compliance data is less flattering. Among accepted papers, 11.9 percent were missing licence information, 4.9 percent lacked a dataset description and 3.5 percent had no URL. These are papers whose entire contribution is a dataset. If the automated report flags a missing licence and the paper is accepted anyway, the report is being read as advice rather than as a requirement. Still, the track produced numbers about its own process that the main track cannot, and that alone makes it the better-instrumented experiment.

What the trajectory implies

If growth continues at anything like the 2020 to 2025 rate, the main track is looking at 30,000 to 40,000 submissions within a few years, and a fixed acceptance rate means 8,000 to 10,000 accepted papers a year at one venue. No poster hall holds that. No reviewer pool of first-time area chairs reviews it well. The two-city format this year was a symptom, and we expect more of them.

The fixes that were tried in 2025 are all about enforcing participation and reducing gaming. None of them reduce the load. The options that do are unpopular: hard caps on submissions per author, desk rejection at scale, a lower acceptance rate, or splitting the conference into subject venues that each carry their own review. The datasets track is the closest thing to a controlled trial of tooling helping reviewers, and its survey suggests tooling helps at the margin. We would like to see the main track instrument itself the same way before the next round of changes, so that in 2027 we are arguing about measured effects rather than about impressions.

Sources

  1. NeurIPS blog: Reflections on the 2025 review process from the program committee chairs (September 30, 2025)
  2. NeurIPS blog: Datasets and Benchmarks track, from art to science in AI evaluations (December 5, 2025)