Agentic engineering patterns: the anti-vibe-coding guide
Simon Willison has started a living guide for professionals who use coding agents, and he is careful to say it is the opposite of vibe coding. Notes on what is in it so far, and why research code needs its own version.
What Willison is building
On February 23 Simon Willison announced a new project, a guide called Agentic Engineering Patterns. The format is deliberate. Rather than a sequence of dated blog posts, it is a set of chapters he intends to keep updating, published one or two at a time each week, with the 1994 Design Patterns book as the model for its structure. He states that every word of the text is his own, with language models used only for supporting work like proofreading and example code.
The definition he settles on is short. Agentic engineering is the practice of developing software with the assistance of coding agents, where a coding agent is a system that runs tools in a loop to achieve a goal, and where the ability to execute code is what separates it from ordinary prompting. The name matters because he is explicitly carving this away from vibe coding, which he keeps in its original sense, prompting a model to write code while you forget the code even exists. Agentic engineering is the other end of that scale, professional engineers using agents to amplify expertise they already have.
The chapters so far
The index already lists more than a dozen chapters across five groups. Under principles there is a chapter arguing that writing code is cheap now, one on hoarding techniques you know how to do so an agent can recombine them, one on using agents to produce better code rather than just more of it, and a chapter of anti-patterns. Under working with agents there are chapters on how the agents work, on using Git with them, and on subagents. Testing and QA gets red/green TDD, a chapter titled First run the tests, and one on agentic manual testing with browser automation. There is also a group on understanding code and a set of annotated real prompts.
The two chapters that were live on launch day tell you where his head is. Writing code is cheap now is about what changes for a team when the cost of a first working draft collapses. Red/green TDD is about getting an agent to write a failing test first, because a test gives the loop something to run against and tends to produce more concise, more reliable code with less prompting. Neither is a trick. Both are old discipline applied to a new tool.
The anti-patterns chapter currently has one entry, and it is the right one. Do not file pull requests with code you have not reviewed yourself. Willison's reasoning is that an unreviewed agent PR shifts the actual work onto whoever has to read it. His fix is to do the first review pass yourself, keep PRs small, check that the description matches the diff, and include evidence of manual testing.
The line that carries the whole guide
The sentence we keep coming back to is from the definitions chapter. Engineers must verify and iterate on the results until they are confident those results address the problem in a credible way. Everything else in the guide is machinery in service of that sentence. Tests exist so the agent can verify. Git exists so you can iterate without fear. Reviewing your own PR exists so the confidence is yours and not the model's.
That framing is why we think the guide will age well even as the tools change. It treats the agent as a fast, tireless, occasionally wrong junior collaborator, and it puts the responsibility for the outcome where it has always been.
Research code has the same problem and no guide
Here is our worry. Everything Willison says about production software applies with more force to experiment code, and nobody has written the equivalent guide. In a research group the agent is generating data loaders, training loops, evaluation scripts and plotting code. Those are the parts of a paper that determine whether a number is real. When generating them was expensive, people reused a small number of scripts they understood. When generating them is cheap, every experiment gets its own fresh, plausible, unreviewed pipeline.
That makes reproducibility worse in a specific way. A bug in a shared evaluation script gets found because many people hit it. A bug in a one-off script that an agent wrote at two in the morning gets found by nobody, and the result it produced goes into a table. The cheapness of generation, which Willison correctly identifies as the central fact for software teams, turns into a cheapness of unverified variation for research teams.
The patterns transfer almost directly. Write the evaluation test before the model code, so the agent has a target that is not the number you hope for. Run the existing tests first. Keep the diff between the baseline pipeline and the new one small enough to read. Never put a result in a paper from a script you have not read yourself, which is the research version of the one anti-pattern he has published.
What we would add for our own use
Two things are missing from the guide as it stands, at least for our purposes, and we would rather write them than wait. The first is a pattern for fixed seeds and pinned environments that the agent is not allowed to change, because an agent that can silently edit the config to make a test pass will do so. The second is a pattern for recording what the agent did alongside the final code, because the trail of prompts and rejected attempts is part of the method now.
We intend to draft both as chapters in the same format and see whether they hold up across a few of our projects. If they do, we will publish them alongside our code. If they do not, that is worth knowing too.
Sources
From the foundation