Auto-GPT and BabyAGI: a post-mortem on the first agent hype cycle
Three weeks after Auto-GPT appeared it is one of the fastest growing repositories on GitHub, and most people who have run it have watched it go in circles. Notes on why the loops spin, what HuggingGPT does differently, and which of the missing pieces look solvable.
What the loops are
Auto-GPT was released on March 30 by Toran Bruce Richards. You give it a name, a role and a handful of goals, and it runs GPT-4 in a loop. Each turn the model proposes a thought, a reasoning, a plan and a command. The command might be a web search, a file write, or a call to a sub-agent. The result comes back into the context and the loop runs again. It has short-term memory for the current task, internet access and can take text and image inputs.
BabyAGI, from Yohei Nakajima, is the same idea stripped to about 140 lines. An execution agent does the current task. A task creation agent reads the result and proposes new tasks. A prioritisation agent reorders the queue. Results go into a vector store so later tasks can retrieve earlier ones. Nakajima is clear that it is a task management loop and that the name is a joke. Both projects went from nothing to tens of thousands of stars in a fortnight.
Why they spin
Run either for an hour and the characteristic failure shows up. The agent completes a subtask, generates three new subtasks that restate the objective in different words, picks one, completes it, and generates three more. The developers attribute the infinite loops to limited context. That is part of it. The model cannot see far enough back to notice it has already done this, and the vector store retrieves whatever is similar to the current task, which is usually the last version of the same task.
The deeper reason is that nothing in the loop knows what done looks like. A goal such as research competitors and write a report has no stopping condition the model can evaluate. So the model, trained to be helpful, keeps being helpful. A human running the same objective would stop when the report existed. The agent has no way to check whether it does, because checking would itself be a task, and the task would generate tasks.
The second failure is confident fabrication. When a search returns nothing useful, the model fills in what it expected to find and moves on, and the error compounds through every later step that depends on it. Wikipedia's summary of the reception captures it. Avram Piltch at Tom's Hardware called it too autonomous to be useful. Clara Shih at Salesforce advised keeping a human in the loop. Both are right about the current version.
The cost nobody budgets for
GPT-4 through the API is priced at roughly 3 cents per thousand input tokens and 6 cents per thousand output tokens. A loop that re-sends its context on every step spends most of that on tokens it has sent before. An agent that spins for an hour can run up a bill with nothing to show, and BabyAGI's own readme warns that running it continuously results in high API usage. Cost is the constraint that ends most experiments, well before the agent finishes.
That has an upside. Cost makes the failure legible. A spinning agent is expensive in a way that a spinning script is not, so people notice and stop it. When the token prices fall, that signal will fade, and the loops will run longer for less reason.
A different design in the same month
HuggingGPT, from Yongliang Shen and colleagues at Zhejiang and Microsoft, was posted the same day Auto-GPT launched, and it is a useful contrast. It also uses ChatGPT as a controller, but the loop is fixed. Task planning breaks a request into subtasks. Model selection picks models from Hugging Face by their descriptions. Execution runs them. Summarisation assembles the answer. Four stages, no open-ended task queue.
The difference is that HuggingGPT knows what done looks like, because the plan is made once and the stages terminate. It gives up generality to get that. It cannot decide halfway through that it needs a different plan. But it does not spin, and the papers we expect to matter over the next year will be the ones that find the middle ground between a fixed pipeline and an open loop.
Which pieces are missing
Three things stand out as absent. Reliable tool calls, so that a command to search or write a file returns structured output the model can check rather than free text it has to parse. Memory that is more than nearest-neighbour retrieval over past results, so an agent can tell it is repeating itself. And cost control built into the loop, a budget that the agent can see and that stops it when spent.
The first looks like a model training problem, and we expect the labs to solve it because it helps every product they sell. The third is an engineering problem and could be solved this week by anyone who wanted to. The second is the one we are least sure about. Knowing that you have already done something is a question about the state of the world rather than the contents of a context window, and we do not see a proposal for it in either repository. If the hype cycle leaves one useful thing behind, we hope it is a clear statement of that problem.
Sources
From the foundation