ICML bans LLM-written text, and the field starts arguing about authorship
ICML's policy for this year prohibits papers whose text was generated by a large language model while permitting light editing. Where the line was drawn, why, and why we think it will not hold.
What the policy says
ICML opens in Honolulu next week and it does so under a rule that did not exist a year ago. Papers that include text generated from a large language model such as ChatGPT are prohibited unless that text is presented as part of the paper's experimental analysis. The program chairs describe this as a conservative position taken because there had been very little time since the November 2022 release of ChatGPT to think through the consequences.
The policy is narrower than the headlines suggested when it was announced. Authors may use a model for light editing of their own text, and the policy compares this to grammar checkers, autocorrect and similar tools. What is prohibited is having the model produce the text. Generated output that appears as an object of study, a sample in a table of model completions for example, is fine. The policy applies to ICML 2023 only and the chairs say it will be revisited.
The questions the chairs were trying to avoid
The rationale the chairs give is a list of open questions rather than a claim about quality. Is generated text novel work or derivative of the training data? Who owns the output of a model? Can generated text constitute plagiarism? Each of these is a real question and none of them had a settled answer in January when submissions were due. Faced with that, the chairs chose to keep generated prose out for a year rather than adjudicate cases they had no framework for.
We have some sympathy for that. A program chair's job is to run a review process, and a review process needs rules that a reviewer can apply. A rule that says "generated text is fine if it is good" pushes the whole problem onto several thousand volunteers with no guidance. A rule that says "no generated text" is at least a rule.
Where the line was drawn, and why it is odd
The line sits between editing and generating, and we think that is the wrong place for it, even though we understand why it was chosen. Consider a non-native English speaker who writes a rough paragraph and asks a model to rewrite it fluently. Is that light editing? The words are mostly the model's. Now consider an author who dictates a detailed outline and asks the model to expand it into prose. The ideas are entirely the author's but every sentence is generated. The policy permits the first case, or seems to, and prohibits the second, and we struggle to see a principled reason.
The deeper issue is that authorship of prose was never what a conference was certifying. A paper is accepted for its claims, its evidence and its method. The prose is the delivery vehicle. Nobody at ICML has ever asked whether an author wrote their own related-work section or had a student draft it. The policy treats the sentences as the thing to protect, when the thing the community actually cares about is whether the results are real and the reasoning is the author's.
Why it cannot be enforced
The policy itself concedes the enforcement problem. There is no automated detection. Violations are investigated only when a reviewer or editor flags a submission, at which point it goes through the same process as a suspected plagiarism case. That means the rule bites only on authors who either admit to it or write badly enough in a recognisable way to get noticed.
Detection tools exist and we have run several of them on our own group's drafts. They flag careful human prose written in a plain register at rates that would make any conference-wide screen unusable, and they are trivially defeated by a second editing pass. The chairs were right not to deploy them. But a rule that cannot be checked becomes, within a year, a rule that is observed only by the conscientious, and a rule that penalises the conscientious is worse than no rule.
There is a second-order effect that we think matters more. The permitted uses, light editing of your own text, are exactly the uses that make papers from outside the anglophone world more readable. If honest authors avoid the tool to stay clearly on the right side of an ambiguous line, the policy will have made the literature slightly harder to read from the very groups it would be least fair to burden.
What we would replace it with
The alternative we would argue for is disclosure and responsibility rather than prohibition. Require authors to state how models were used, in a short section like the compute or ethics statements the community already accepts. Then hold authors fully responsible for every sentence regardless of how it was produced. A hallucinated citation is the author's hallucinated citation. A fabricated result is the author's fabrication. That rule is checkable, it is fair to non-native writers, and it puts the weight on the thing conferences actually certify.
ICML said this policy is for 2023 only. Our prediction is that next year's version, from this conference or one of its neighbours, will look much closer to disclosure than to a ban, and that within two years the question of whether a model touched the prose will seem as quaint as asking whether a spellchecker did. What will not go away is the question of whether the results are real. That one deserves all the attention the prose question is currently taking.
Sources
From the foundation