Swarm and AutoGen: the multi-agent framework question
OpenAI's Swarm is a small library built on two primitives, agents and handoffs, and is labelled educational. Microsoft's AutoGen is a conversation-programming framework with a 43-page paper behind it. A comparison, and a question about whether multi-agent orchestration is an abstraction or a workaround.
Two libraries, two philosophies
OpenAI's solutions team has published Swarm, a small Python library it describes as an educational framework for exploring ergonomic, lightweight multi-agent orchestration. The repository is explicit that it is experimental, is not intended for production, and is managed by the solutions team rather than the API team. The accompanying cookbook entry, Orchestrating Agents: Routines and Handoffs, says the library is a sample only and encourages developers to adapt the ideas rather than import the code.
AutoGen, from Qingyun Wu and thirteen collaborators at Microsoft and partner institutions, is the opposite kind of artefact. The paper was posted in August 2023, revised in October, and runs to 43 pages including appendices. It introduces conversable agents, entities that can combine LLMs, human input and tools, and conversation programming, a way of defining how those agents interact using natural language and code together. It is evaluated on mathematics, coding, question answering, operations research, online decision-making and a game. The framework is positioned as generic infrastructure for applications of various complexities.
What Swarm actually is
Swarm has two primitives. An Agent is a set of instructions plus a list of tools. A handoff is what happens when a tool call returns another Agent instead of a string. The runtime notices the return type, swaps the active agent, and continues the conversation with the full history intact. That is the whole mechanism. The cookbook builds it up from a routine, which it defines as a list of natural-language instructions represented as a system prompt along with the tools needed to complete them, and then adds the transfer function, something like transfer_to_refunds that returns the refunds agent. A Result object lets a function return a value, switch agents and update context variables in one step.
The library is stateless by design. It runs almost entirely on the client, keeps nothing between calls, and mirrors the Chat Completions API in that the caller must pass prior messages and context into each run. That makes it easy to test and easy to reason about, and it also means the library takes no position on memory, persistence, retries or parallelism. Those are left to the user, which is a reasonable choice for a teaching example and a limiting one for anything else.
What AutoGen adds and what it costs
AutoGen's core idea is that a multi-agent application is a conversation, and that the programmer's job is to define who speaks to whom, under what conditions, and what each participant is allowed to do. Agents can be backed by a model, by a human in the loop, by a tool executor, or by combinations. Control flow can be expressed in natural language, with agents deciding when to reply or hand over, or in code, with explicit orderings. The paper's applications show the range, from a two-agent math solver to an entertainment application.
The cost of that generality is that the abstraction is larger than the problem for most uses. When the conversation is the control flow, debugging means reading transcripts, and the failure modes are the ones anyone who has watched two models talk to each other will recognise: loops, mutual deference, and long exchanges that converge on nothing. Swarm's handoff model avoids this by making the transfer a deterministic event triggered by a function call. Only one agent is active at a time, and the graph of who can hand to whom is fixed in code.
Abstraction or workaround
The question we keep coming back to is why anyone wants multiple agents at all. The honest answer, in most of the deployments we have seen, is context. A single model with one long system prompt covering triage, sales, refunds and escalation performs worse than four models with short focused prompts, because the long prompt has instructions that conflict, tools that distract, and a context window filling with material irrelevant to the current turn. Splitting the work into agents is a way of managing what the model sees. The handoff is a context switch.
If that is the reason, then multi-agent orchestration is a workaround for a limitation of current models rather than a durable abstraction, and it should get less necessary as models get better at ignoring irrelevant instructions and handling long contexts. Swarm's own design suggests its authors think so. A framework that fits in a cookbook, is labelled educational, and whose main construct is a function that returns a different prompt, is not staking a claim that agents are a new kind of software component. AutoGen's framing, in which the conversation is the program, makes the stronger claim, and it is the one we are less sure survives the next two model generations.
There is a case for the stronger view. Separate agents can run on separate models, with separate permissions, separate tools and separate audit trails, and those are properties of a system rather than of a context window. If the reason to split is security or cost rather than prompt hygiene, the abstraction earns its keep regardless of how good the models get. We do not think the frameworks distinguish between those reasons yet, and we would like one that did.
What we would try
Take a task that is currently being solved with a handoff graph, collapse it into a single agent with the union of the tools and a merged prompt, and measure the gap. Then repeat with each new model release. If the gap closes, the multi-agent version was a context workaround and can be retired when the model catches up. If it does not, something about the decomposition is doing real work, and that is the thing worth naming and building a framework around. Swarm is a good scaffold for the experiment precisely because there is so little of it.
Sources
From the foundation