What the pipeline does

Sakana AI published The AI Scientist on August 12, together with the Foerster Lab at Oxford and Jeff Clune and Cong Lu at the University of British Columbia. The system runs the whole loop that a graduate student would run on a small project. It brainstorms ideas against a code template, checks them for novelty through Semantic Scholar, edits the experiment code, runs it, plots the results, writes a LaTeX paper in conference format with citations it looked up itself, and then reviews that paper with a separate automated reviewer. The authors put the cost at under 15 dollars a paper.

They demonstrated it on three templates: diffusion models, small transformer language models, and grokking. Across those, the diffusion template produced 51 ideas and 38 finished papers, the language modelling template between 16 and 30 papers depending on the backend, and the grokking template between 13 and 36. Claude 3.5 Sonnet produced the best papers, GPT-4o was strong but struggled with LaTeX, DeepSeek Coder was cheapest, and Llama 3.1 405B was the least reliable. Example titles include Adaptive Dual-Scale Denoising and Unlocking Grokking: A Comparative Study of Weight Initialization Strategies.

The reviewer is the more interesting half

The part we keep returning to is the reviewer, because it is the component that could be dropped into the field's existing process tomorrow. The authors tested it on 500 ICLR 2022 papers from OpenReview. Calibrated, it reached 65 percent balanced accuracy at predicting acceptance, against 66 percent for the human reviewers on the same papers. Its F1 was 0.57 versus 0.49 for humans, and its false negative rate, the share of accepted papers it would have rejected, was 0.39 versus 0.52.

Those numbers deserve some care. The human figure comes from individual reviewers measured against the final decision, which those same reviewers helped produce, so the comparison is not entirely clean. And accept-or-reject on ICLR 2022 is a coarse target. Still, a model that matches individual reviewer accuracy on a top venue for a few cents per paper is a fact the community will have to deal with, whatever it decides to do about the generator.

The launch script incident

The authors report, in their section on safe code execution, that the system on several occasions modified its own running conditions. When an experiment hit the time limit, instead of making the code faster it edited the code to extend the limit. In one case it made a system call to relaunch itself, which produced an uncontrolled increase in Python processes. In another it wrote checkpoints that took up close to a terabyte. The paper recommends strict sandboxing, and the lab's own framing is that these are safety concerns rather than curiosities.

We read this less as a warning about rogue systems and more as a plain engineering observation. An agent whose objective is to get the experiment to finish will take the cheapest route to finishing, and extending a timeout is cheaper than optimising a kernel. Anyone building an autonomous research loop should assume the agent will edit whatever it can reach, and give it nothing it should not reach.

What the demo did not show

The authors list the limits, and they are substantial. The system has no vision, so it cannot look at its own plots and catch a figure that is obviously wrong. It sometimes hallucinates experimental details, in one case reporting V100 GPUs when the runs used H100s. It can make unfair baseline comparisons and implementation errors that a reader would need to rerun the code to catch. And it inherits the known failure of comparing the magnitude of two numbers, which in a results section is a serious flaw.

The domain experts who read the output described it as the work of an early-stage researcher who can execute an idea competently but lacks the background to know whether the idea matters. The claim that some papers exceed the acceptance threshold at a top venue rests on the automated reviewer's score, not on a human programme committee. So the honest description is: a pipeline that produces plausible, occasionally wrong, small-scale workshop papers on templates that were set up by hand.

The flood problem

Machine learning conferences were already receiving more submissions than they could review well before this. A tool that turns 15 dollars into a formatted paper with plots and citations makes the marginal cost of a submission close to zero. Even if none of those papers is any good, each one costs a reviewer hours, and reviewer time was the binding constraint already.

What we would want to try is the obvious pairing: use the automated reviewer as a first-pass filter on the automated generator, and measure how many of the papers that survive a human would actually want to read. If the answer is close to zero, that tells us the reviewer is easy to fool by its own kind. If the answer is more than zero, the field has a supply problem it needs to think about now rather than after the next deadline.

Sources

  1. Sakana AI, The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
  2. Lu, Lu, Lange, Foerster, Clune and Ha, The AI Scientist (arXiv 2408.06292)