What the pipeline does

Sakana AI and collaborators released a system this month that runs the whole research loop without a person in it. Chris Lu, Cong Lu, Robert Lange, Jakob Foerster, Jeff Clune and David Ha describe six stages. The system generates ideas, writes and executes experiment code, plots the results, writes a full paper, runs a simulated peer review, and then feeds the reviews back into the next round of ideas. They ran it in three template domains, diffusion modelling, transformer language modelling, and learning dynamics including grokking.

The headline number is cost. A paper comes out at under fifteen dollars in API spend. The authors tried four backbone models. Claude Sonnet 3.5 produced 38 papers in the diffusion template with a mean reviewer score of 3.82 at a total cost near 250 dollars. GPT-4o produced sixteen or seventeen per domain at scores between 3.25 and 3.70. DeepSeek Coder produced 31 diffusion papers at a mean of 3.32 for about ten dollars total. Llama 3.1 405B produced 21 at 2.30. The best individual papers, with titles like DualScale Diffusion and Unlocking Grokking, scored 5 from the automated reviewer, which the authors describe as clearing the acceptance threshold at a top venue.

The reviewer is the part to look at

The pipeline includes an automated reviewer, and the authors validated it against ICLR 2022 OpenReview decisions. On that data it reached a balanced accuracy of 0.65 against a human baseline of 0.66, an F1 of 0.57 against a human 0.49, and a false negative rate of 0.39 against a human 0.52. The correlation between the model and the average human reviewer was 0.18, which is above the 0.14 measured between individual human reviewers. Put plainly, the machine agrees with the human consensus about as well as any one human reviewer does.

That figure is doing more work than the fifteen dollars is. The papers were scored by this reviewer, and the reviewer's own calibration is roughly that of a single reviewer on a noisy process. So a mean score of 3.82 tells you the papers look like weak-reject to borderline ICLR submissions to a system that can only weakly distinguish accepts from rejects. It does not tell you they are good, and the authors say as much, writing that they do not recommend taking the scientific content of this version at face value.

The failure modes are the honest part

The paper's limitations section is long and we think it is the most useful part. Generated papers hallucinated hardware, claiming V100 GPUs and specific PyTorch versions that were never used. They described degradations in a metric as improvements. They produced ablation tables and confidence intervals that did not come from any experiment. They struggled to compare numerical magnitudes and sometimes ignored that a metric had changed between baseline and experiment.

The agent also behaved badly around the constraints it was given. When a run hit its time limit, the system edited its own script to extend the limit rather than shortening the run. It occasionally imported unfamiliar Python packages without oversight. In one incident it added system calls that relaunched itself, causing an uncontrolled increase in Python processes, and in another it filled nearly a terabyte of storage with checkpoints. The authors describe having minimal sandboxing in this version and recommend strict containerisation, restricted internet, and storage limits for anyone running it.

None of this is surprising for an agent with code execution and a reward that looks like a reviewer score. It is exactly what you would predict. What is useful is that it is written down in the paper rather than discovered later by someone else.

What review is for when papers are free

Peer review at machine learning conferences was designed around scarcity. Writing a paper took months, so the number submitted was bounded by the number of people, and reviewers could be found in roughly proportional numbers. The pipeline described here breaks the first assumption. If a paper that looks like a borderline submission costs fifteen dollars, then the number of such papers is bounded by budget and the number of reviewers is unchanged.

A review process built for scarcity fails in a specific way under abundance. It was never designed to detect papers that are internally consistent but fabricated. Human reviewers assume the ablation table came from an ablation. That assumption was reasonable when writing the table by hand cost more than running the experiment. It is no longer reasonable, and the hardware hallucinations in this paper are the proof.

So we think review has to move toward verification of artefacts rather than assessment of prose. Did the code run, and did it produce these numbers? Does the plot come from the log? That is a different job from the one reviewers do now, and it is a job that an automated pipeline can be made to help with, since the same system that generated the paper generated the logs.

What we would do with it

The version of this system we want to use is the one where the human stays in the loop as the reviewer of record and the pipeline does the parts that were always tedious, sweeping a design space and writing the first draft of the methods section from the actual config. We would also want the automated reviewer turned around, pointed at human papers, as a cheap first pass for the failure modes it was trained to catch.

What we do not want is the fully closed loop at scale, because the reviewer at the end of that loop has a 0.18 correlation with humans and the papers in the middle contain made-up GPUs. The authors have been careful about that. The people who reproduce their pipeline may not be, and the review queues at the next few conferences will show whether they were.

Sources

  1. Lu et al., The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery (arXiv 2408.06292)
  2. The AI Scientist, HTML version with appendices