Building effective agents: the essay that told everyone to use fewer frameworks
Anthropic's December essay separates workflows, where code decides the path, from agents, where the model does, and argues for starting with direct API calls and five composable patterns before reaching for a framework. Reading notes from a year in which frameworks multiplied faster than working agents.
A definition that removes most of the hype
The essay Anthropic published on December 19, credited to Erik S. and Barry Zhang, does its most useful work in the first few paragraphs, before any pattern is described. It splits what people have been calling agents into two kinds of system. A workflow is one where LLM calls and tools are orchestrated through predefined code paths. An agent is one where the model dynamically directs its own process and tool use and stays in control of how the task gets done. Most of what shipped this year under the agent label is, by this definition, a workflow, and the authors say that plainly.
The reason we find the split useful is that it separates two questions that had been blurred together. Whether a system is reliable depends mostly on how much of the control flow is written in code. Whether a system is flexible depends mostly on how much is left to the model. Once the two are named, the design question becomes how much control to hand over for a given task, which is a tractable engineering decision rather than a matter of belief about what agents can do.
The building block and the five patterns
Everything in the essay is built from one component the authors call the augmented LLM, a model with retrieval, tools and memory attached, that can generate its own search queries, choose which tool to call and decide what to keep. The workflows are then arrangements of that block. Prompt chaining runs steps in sequence with each call working on the previous output. Routing classifies an input and sends it to a specialised prompt. Parallelization either splits a task into independent pieces or runs the same task several times and votes.
The two remaining patterns hand more to the model. In orchestrator-workers, a central call breaks a task down at runtime and delegates the pieces, which differs from parallelization in that the subtasks are not known in advance. In evaluator-optimizer, one call produces a draft and another critiques it in a loop, which the authors present as worthwhile when there are clear evaluation criteria and iteration measurably helps. Reading the list, we noticed that four of the five patterns can be drawn as a diagram with the control flow entirely in code. Only the orchestrator is close to an agent, and even it is bounded.
The advice on frameworks
This is the paragraph that got the essay passed around. The authors acknowledge that frameworks make it easy to get started, then say that they add layers of abstraction which obscure the underlying prompts and responses, make debugging harder, and tempt people into complexity they do not need. Their recommendation is to start by using the model API directly, since many of the patterns above take only a few lines of code, and if a framework is used, to make sure you understand what it is doing underneath.
We read that against the year we have had. Since the spring there has been a new agent framework roughly every few weeks, each with its own abstractions for chains, tools, memory and planning, and each one making it slightly harder to see what string was actually sent to the model. Teams we have talked to spent more time fighting the abstraction than the task. The essay's position is that the abstraction was never carrying much, because the patterns are simple enough to write by hand, and that the cost of not seeing your prompts is higher than the cost of writing a loop.
When an agent is actually the right call
The essay is not against agents, and the section on when to use them is careful. Agents fit open ended problems where the number of steps cannot be predicted and a fixed path cannot be hardcoded, and they make sense in environments where the model's actions can be trusted at scale. The costs named are higher spend and the compounding of errors over many steps, and the implicit advice is that trust has to be earned before autonomy is handed over.
The two appendix examples make the case for where those conditions hold. Customer support has a conversational shape, tools for looking up orders and issuing refunds, and a clear signal of success. Coding is the stronger example, because a code change can be checked against automated tests, the agent can iterate on the failures, and the output is objectively verifiable. That pattern, an open task with a cheap, reliable verifier at the end, is the one place where we expect fully agentic systems to hold up in the next year, and it is no accident that coding is the example the authors chose to expand.
Three principles, and what we would add
The closing principles are simplicity, transparency through explicit planning steps, and careful design of the agent computer interface with thorough tool documentation and testing. The last is the one we think teams underinvest in. The tool definition is a prompt. The examples, parameter names and error messages shape what the model does as much as the system prompt, and a tool that is clear to a human engineer is not automatically clear to a model.
What we would add is measurement. The essay tells you to start simple and add complexity only when it demonstrably improves outcomes, which presupposes an evaluation that can show the improvement. Most teams we know do not have one, and without it the argument for simplicity becomes a matter of taste rather than evidence. The right next step for anyone taking the essay seriously is to build the twenty task evaluation for their use case before building the agent, and to keep the workflow that scores best rather than the one that looks most autonomous.
Sources
From the foundation