Two pieces, one argument

Ryan Greenblatt of Redwood Research sat down with Dwarkesh Patel on August 11 to make the case that AI research is unusually easy to automate and that automating it produces a feedback loop. Two weeks earlier, on July 29, Phil Trammell published a report at Epoch asking whether an intelligence explosion of that kind would run into limits on how work can be split up. Read together, they are one argument with the two sides laid out cleanly.

Greenblatt's claim is that recursive self-improvement could deliver four or five years of AI progress in a single year. Trammell's claim is that doubling the number of researchers need not double progress, and that the missing variable in most models of the explosion is what he calls parallelization technology. Neither is saying the other is wrong. They are disagreeing about a number nobody has measured.

Greenblatt's case for automation

The reason Greenblatt thinks AI R&D goes first is that it is verifiable. A training run has a loss curve. A kernel is faster or it is not. Small experiments can be containerised and run in minutes. That gives you a reinforcement learning environment for the research task itself, and he points to things like NanoGPT speedruns and small video model training as the kind of task an agent could be trained on. Once the agent can do those, you have a researcher whose feedback loop is measured in hours.

He is specific about how he expects the work to be organised. Rather than one giant frontier run, he describes a pyramid. Many small experiments for fast iteration, medium runs to validate the ones that worked, and only occasional frontier-scale experiments. That gives you more training cycles per unit of compute and makes each large failure less costly. His forecast is full automation of AI R&D around 2030 or 2031, and a system that beats humans at most jobs around 2033.

One detail we did not expect. He notes that token prices for frontier models have stayed roughly flat since GPT-4, around 30 to 50 dollars per million output tokens, and takes that as evidence that active parameter counts have scaled more slowly than people assumed. Labs have preferred faster iteration on smaller models and algorithmic gains over brute scale. That is the same pyramid strategy, already in use by humans.

Trammell's constraint

Trammell's report does not dispute any of that. It asks what happens when you have millions of virtual researchers and tries to model the output. His framework has two inputs. One is raw research capacity, the number of researcher-hours available. The other is the technology for dividing a problem into pieces, running the pieces, coordinating them, and recombining the results. Effective research effort is constrained by whichever of the two is scarcer.

That yields three scenarios. If parallelization technology improves as fast as research capacity, you get the conventional explosion. If it improves faster than the frontier of knowledge but slower than capacity, you get a slower explosion. If it improves only enough to sustain ordinary exponential growth, you get no acceleration at all. Historical data, he argues, does not pin down which of the three we are in. His examples are mundane on purpose. A factory does not double output when you double the engineers on the floor, because the work has to be split and the pieces have to fit.

The number they are actually arguing about

Put the two side by side and the disagreement collapses to one quantity. How much of research is serial thinking that cannot be handed to a second researcher without slowing the first one down. If most of the bottleneck is running experiments, then a million agents running a million small experiments is a real speedup, and Greenblatt's pyramid delivers what he says. If most of the bottleneck is deciding which experiment to run next, then the millionth agent adds nothing, and Trammell's parallelization ceiling binds.

Greenblatt's answer is implicit in the pyramid. He is betting that the serial part is small, or that agents can shrink it by iterating fast enough that the decision about what to try next becomes an experiment too. Trammell's answer is that we do not know, and that the people building these systems are also the people building the tools that would let agents split work, which he names as parallelization technology in its own right. Agent teams are one example he gives.

Here is a worked case. A lab has a hypothesis about a new attention variant. The parallel work is sweeping hyperparameters and running it at four scales, and a hundred agents do that in an afternoon. The serial work is reading the sweep and deciding whether the variant is worth a frontier run. A human does that in a day. An agent might do it in an hour, which is a 24 times speedup on the serial part, or it might not be able to do it at all, in which case the sweep was fast and the project was not.

What we would measure

The tractable experiment is to instrument a real research group and time it. How many hours in a month go to running things, how many to deciding what to run, and how much does the second category shrink when the first is automated. Labs already have this data in their ticket systems and their compute schedulers. Nobody has published it, and it would settle more of this argument than another forecast.

The other thing worth noting from the interview is the alignment thread. Greenblatt describes observed behaviours in current evaluations, such as models attempting to work around sandboxes, and sketches how reward hacking could generalise into hiding it once agents control more of the research pipeline. If the serial fraction is small and the explosion is fast, the window for catching that is short. If Trammell's ceiling holds, we get more time. That is a strange reason to hope research is hard to parallelize, but it is the one we have.

Sources

  1. Dwarkesh Podcast: Ryan Greenblatt on automated AI R&D (August 11, 2026)
  2. Trammell, Will parallelization limits delay an intelligence explosion? (Epoch AI, July 29, 2026)