NeurIPS desk-rejects 18 percent of position papers for being written by AI
The position paper track ran every submission through a detector, rejected 178 without appeal and asked 123 more for version history. Notes on the numbers, the appeal, and whether a detector should be the thing that decides.
What the track did
The NeurIPS 2026 position paper track told authors in advance that papers had to be substantially human-written, with AI limited to copy-editing and similar peripheral changes. On June 2 the organisers published what happened when they checked. They ran all 969 submissions through Pangram, version 3.3.2, under a zero data retention agreement, and 273 of them, 28.2 percent, scored 100 percent on the detector's AI metric. Another way to cut the same data is that 42.7 percent scored 90 or above and 70.5 percent scored 50 or above.
The comparison row is the part that made us sit up. The same track last year had 8.2 percent of submissions at a perfect score. The NeurIPS ethics and society track this year had 2.1 percent. FAccT 2025 had none. So this is a jump of more than three times within one venue in one year, in a track whose whole point is argument rather than results.
Who was rejected and who was asked
The enforcement had tiers. Papers scoring 0.9 or higher were desk-rejected, 77 of them. Papers between 0.8 and 0.9 were desk-rejected when an additional factor was present, such as multiple solo-authored submissions or the author's history, which caught another 79. And 22 papers scoring 0.5 or above were rejected because the authors had not disclosed AI use. That is 178 papers, 18.4 percent of the track, rejected without appeal.
A further 123 papers, 12.7 percent, scored between 0.8 and 0.9 with no aggravating factor and were given until June 15 to make a case. The required evidence is specific. An online version with edit history showing a checkpoint before AI was used and one after. An analysis showing the AI edits added no substantive content beyond the pre-AI version. Evidence that a human verified the text after the edits. In effect, the appeal asks authors to prove a negative with a document trail they may not have kept.
How much the window size matters
The detail that should be read before anyone forms a view is in the methodology. Pangram scores a paper by dividing it into windows of 250 to 350 words by default, and a window counts toward the AI score if its probability crosses 0.75. The organisers reran the analysis with windows of roughly 100 words. The share of papers scoring 100 percent fell from 28.2 percent to 2.16 percent.
The organisers present this as a refinement that let them separate heavy AI generation from lighter use, and the enforcement thresholds were set on the refined analysis. But it is worth stating plainly what that sensitivity means. The headline number in the post's title depends on a windowing parameter, and a different reasonable choice would have produced a headline of two percent instead of twenty-eight. Anyone who cites the 28 percent figure without the 2.16 percent figure beside it is not describing the report.
The calibration they ran
To their credit, the organisers did not take the detector on faith. They wrote twelve scenarios covering the range from proofreading to full generation and checked what Pangram made of each. Proofreading and light copy-editing were never flagged. Heavy line editing, structural rewriting, hybrid revision and back-translation sat in a borderline band. Generation from a one-sentence plan, substantive rewriting by the model, original model-authored passages and human edits layered on top of model text were flagged consistently. In a separate test, completions in which the model wrote 20 percent or less of the text were never flagged.
This is the right kind of experiment and it is also a small one. Twelve scenarios is not a distribution, and the scenarios were written by the people setting the policy, who know what the detector is sensitive to. The result tells us the detector behaves sensibly on cases its operators can imagine. It does not tell us the false positive rate on a non-native English speaker who writes in short declarative sentences, which is the case the detector-as-gatekeeper debate has always turned on.
What substantially human-written should mean
A position paper is an argument. The thing being evaluated is whether a person holds a view, has thought it through and can defend it. That is a different object from an empirical paper, where the contribution is the experiment and the prose is a wrapper. If a model wrote the argument, there may be no author who believes it, and that is a legitimate reason for this track to be stricter than the main conference. The policy is defensible on its own terms.
The mechanism is where we part ways. A detector score is a statistic about word choice, and this report shows it swinging by an order of magnitude with a preprocessing setting. Using it to triage which papers get a closer look is reasonable. Using it to reject without appeal at a fixed threshold makes the threshold the policy, and the threshold was chosen after seeing the distribution. The version-history requirement points at a better design. If a track wants human-written arguments, it could require the edit history up front, from everyone, and read it only when something looks off. That shifts the burden from an unreliable classifier to a record the author controls, and it would also have given the 123 authors now scrambling for evidence something they already had.
Sources
From the foundation