Estimate the corpus, not the review

Weixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp and eight coauthors at Stanford posted a paper on March 11 that sidesteps the problem everyone else has been failing at. Per-document AI detectors are unreliable and easy to evade, and accusing an individual reviewer on the strength of one is indefensible. So the authors do not classify documents. They estimate a single number, the fraction of a corpus that was substantially modified by a language model, using a maximum likelihood fit over word frequencies.

The mechanics are simple enough to check. Take a set of human-written reference texts and a set of AI-generated ones, and estimate a probability distribution over words for each. A mixed corpus is then modelled as a mixture of the two distributions with an unknown mixing weight, and that weight is fit by maximum likelihood. The trick is choosing a vocabulary, mostly adjectives and adverbs, where the two distributions differ most. On held-out validation sets the method's error is under 2.4 percent, and under 1.8 percent on ICLR 2023 reviews.

What the numbers say

Applied to reviews submitted after ChatGPT's release the estimates are 10.6 percent for ICLR 2024, across 27,992 official reviews, 9.1 percent for NeurIPS 2023, 6.5 percent for CoRL 2023 and 16.9 percent for EMNLP 2023. The Nature portfolio journals, used as a control, show no significant increase. The word-level evidence behind the estimates is striking on its own. In ICLR 2024 reviews the word commendable appears 9.8 times as often as before, intricate 11.2 times, and meticulous 34.7 times.

It is important to read the headline correctly. Substantially modified does not mean fully written by a model. A reviewer who wrote their own review and asked a model to polish it counts. A reviewer who pasted the paper in and asked for a review also counts. The method cannot separate those, and we think the 17 percent figure has already been reported in places as though it could.

Where the usage clusters

The part of the paper that should worry program chairs is the correlational analysis. Reviews submitted within three days of the deadline show higher estimated LLM usage. Reviews where the reviewer rated their own confidence at two or below on the five-point scale show higher usage. Reviews that contain a scholarly citation, detectable by the presence of et al., show lower usage. Reviewers who reply less to author rebuttals show higher usage. And reviews that sit close to the centroid of the corpus in embedding space, meaning the ones that say what every other review says, show higher usage.

Put together, that is a portrait of the disengaged reviewer. Late, unsure, uncited, absent from the discussion, and generic. None of that is new. What is new is that a tool now makes the disengaged review look like a competent one, and the signals a chair used to rely on to spot a thin review are being smoothed away by the same tool.

Limits of the method

The estimate depends on the reference distributions. The AI-generated reference texts were produced by a particular model with particular prompts, and the method is calibrated to that. A different model, or an instruction to avoid the telltale adjectives, would shift the fit in ways the validation sets cannot catch. The authors are open about this, and it is the reason the number is a floor as much as an estimate. It is also a corpus-level tool by construction. It tells a chair that a tenth of the reviews are affected. It cannot tell them which tenth, and nobody should pretend it can.

What a conference can do

The obvious lever is policy, and most venues now have one, but a policy without measurement is a hope. Liang's method gives a chair a cheap way to measure the corpus every cycle and report the number publicly, which is at least a deterrent and at most a way to check whether a policy change did anything. We would want that number in every program chair report from now on, alongside acceptance rates.

The less obvious lever is the correlations. If usage clusters in late, low-confidence reviews, then the fix is structural. Earlier deadlines with real enforcement, fewer papers per reviewer, and a norm that a confidence of two is a signal to reassign rather than to accept a padded review. The model is filling a gap that overloaded reviewing created. Measuring the fill is useful. Closing the gap is the actual work, and we would like to see a venue try it and report the number before and after.

Sources

  1. Liang et al., Monitoring AI-Modified Content at Scale (arXiv 2403.07183)