Two launches, one shape

On February 24 Anthropic announced Claude 3.7 Sonnet and, in the same post, a command-line tool called Claude Code, described as its first agentic coding tool and released as a limited research preview. The tool searches and reads code, edits files, writes and runs tests, commits and pushes to GitHub, and calls other command-line programs, with the developer approving actions as it goes. On April 16 OpenAI released Codex CLI, an open-source coding agent under Apache 2.0 that runs locally in the terminal and does the same set of things.

What strikes us is how little the two disagree about form. Neither is an editor plugin. Neither is a chat window with a code block. Both are programs you start in a directory that then read and change that directory, and both put the human in a loop of approvals rather than in a loop of copying and pasting.

Why the terminal

The terminal was already the interface through which every other tool in a software project is driven. Tests, linters, build systems, package managers and version control all expose themselves as commands that take text and return text with an exit code. An agent that lives in the shell gets all of them for free, and gets them in the form a language model handles best, which is text. An agent that lives in an editor has to be given each of those capabilities through an integration, and the integrations are where the effort goes.

The second reason is the feedback loop. A coding agent is only as good as its ability to check its own work, and in the terminal the check is a command. Run the tests, read the failure, edit, run again. Claude Code's documentation now describes exactly that loop, and adds that the tool composes with pipes and can be run non-interactively in CI. Git is the third reason. Commits give the agent a unit of work with a natural checkpoint and give the human a diff to review, which is a better review surface than watching an editor buffer change.

What the two tools do differently

The main difference at launch is openness. Codex CLI is open source and the repository is public, so anyone can read how it decides when to ask permission and how it sandboxes commands. Claude Code is a closed binary distributed through a package manager, and its permission model is documented rather than inspectable. For a research group this matters, because we want to study how the agent loop is built as well as what it produces.

The other difference is how much of the product is the model. Anthropic launched Claude Code alongside a model it reports as state of the art on SWE-bench Verified and TAU-bench, and the tool is a way to get that model into a repository. OpenAI's tool ships as a scaffold that can be pointed at its API models, and the tool itself is the object being released. Whether the value sits in the model or in the loop around it is an open question that these two launches answer differently.

What the research preview label hides

Both launches carry a hedge. Anthropic calls Claude Code a research preview and says it wants developer feedback. The label is honest as far as it goes, and it also quietly sets expectations about two things that are not in the announcement. The first is reliability. An agent that can run arbitrary commands and push to a remote can also delete the wrong directory or commit a secret, and the approval prompts are the only thing standing between a confident model and an irreversible action. Neither announcement gives a failure rate or a description of what went wrong in internal use.

The second is cost. Claude 3.7 Sonnet is priced at 3 dollars per million input tokens and 15 dollars per million output, with thinking tokens billed at the output rate. An agent that reads a large repository re-sends much of that context on every turn. Nothing in the launch post tells a user what a typical session costs, and from our early runs the answer varies by more than an order of magnitude depending on how much of the codebase the agent decides to read. A preview label is a fair way to say the product is early. It is not a substitute for those two numbers.

What we are going to measure

We are setting up both tools on the same set of tasks, small bug fixes in repositories we know well, with the approval mode set to ask on every command so we can log what the agent wanted to do alongside what it did. The questions are how often each proposes a destructive action, how many tokens a completed task consumes, and how often the tests pass on the first run against the third. Those three numbers are the ones we would want on a launch page, and the terminal makes them easy to collect, which is itself an argument for the design.

Sources

  1. Anthropic, Claude 3.7 Sonnet and Claude Code
  2. OpenAI, openai/codex on GitHub
  3. Wikipedia, OpenAI Codex (AI agent)
  4. Claude Code documentation, Overview