What happened

On March 12 Sakana AI announced that a paper produced by its AI Scientist-v2 system had passed peer review at an ICLR 2025 workshop. The workshop was ICBINB, which as it happens is a venue for negative and surprising results. Three manuscripts were submitted. One received reviewer scores of 6, 7 and 6, an average of 6.33, which the company says is above the average threshold for acceptance at that workshop. The other two did not pass.

The experiment was run with the cooperation of ICLR leadership and the workshop organisers, and had ethics approval from the University of British Columbia. Reviewers were told that AI-generated papers were in the pool but were not told which ones. Human involvement was limited to choosing the broad research topic and picking which three of the generated papers to submit. The text, the experiments, the figures and the citations were produced by the system without edits.

Then the accepted paper was withdrawn before publication. Sakana's stated reason is that the AI and scientific communities have not decided whether they want AI-generated manuscripts in the literature, and the company did not want to make that decision unilaterally by letting it stand.

How much the score means

Not much, and the announcement says so. A workshop is a lower bar than the main conference, and this one accepts something like 60 to 70 percent of submissions. One paper out of three cleared it. Sakana also says the team's own read was that none of the three met their internal standard for a main-track ICLR paper. So the headline is that a system produced a manuscript that three workshop reviewers thought was fine, and produced two that they did not.

That is still a real result. Two years ago the question was whether a language model could write a coherent methods section. Now the question is whether a fully automated pipeline can get through a review process designed for humans, and the answer is sometimes. What it says about the quality of the science inside the paper is a separate question that a review score does not answer, and the company is right to keep those apart.

The absence of a norm

What we find more interesting is the withdrawal. The paper passed. By the rules of the venue it was accepted. There was no rule that said what to do with an accepted paper that had no human author, because nobody had written one. So the company and the organisers agreed in advance that if a paper got in, it would come out again. The result was a review process that reached a verdict and then discarded it, by mutual consent, because the verdict had no category to land in.

Consider what the alternatives would have been. Publish it, with the system listed as author, and the workshop has set a precedent it did not intend to set. Publish it under a human name and the experiment becomes a deception. Reject it after acceptance and the reviewers' judgement is overruled for reasons unrelated to the paper. Withdrawal was the only option that did not commit anyone to a position, and that is the whole point. The community has no position.

Here is a concrete version of the problem. A graduate student uses a system like this to generate a draft, reads it, agrees with it, runs a couple of the experiments to confirm, and submits it under their own name. Is that a human-authored paper with AI assistance, or an AI paper with a human signature? Every venue's policy on AI use was written with editing and translation in mind. None that we have read anticipates a pipeline that picks the hypothesis.

What reviewers were actually testing

There is a second reading of the result that we think is underrated. Peer review at a workshop is a filter for whether a manuscript looks like a paper, cites the right things, and makes a claim its evidence supports. A system optimised on a corpus of accepted papers will get good at looking like a paper. The score tells us that the filter can be passed by a system that has learned the form. It does not tell us the filter was ever measuring the thing we cared about.

That is a finding about review as much as about the system. If a workshop's acceptance criteria can be met by a pipeline with no human in it, then the criteria are checking surface features, and the human papers that pass are being checked for the same surface features. The ICBINB organisers deserve credit for agreeing to a test that risks that conclusion about their own process.

What we would want written down

The useful next step is for a major venue to publish a policy, before the next cycle, that says what happens to a fully generated submission. Any answer is better than the current silence. Disclose and review as normal, disclose and route to a separate track, or reject outright. We would prefer the second, because the interesting science question is what these systems produce when they are allowed to compete, and the interesting policy question is what the community does with a track full of them.

In the meantime, a withdrawn paper that passed review is the most accurate description available of where things stand. The tools ran ahead of the rules by exactly one workshop.

Sources

  1. Sakana AI: The AI Scientist generates its first peer-reviewed scientific publication