The compiler

Nicholas Carlini's post describes sixteen Claude Opus 4.6 agents working for nearly two weeks on a C compiler. The result is about 100,000 lines of code, produced across almost 2,000 Claude Code sessions, consuming 2 billion input tokens and 140 million output tokens, at a total cost just under 20,000 dollars. It compiles the Linux 6.9 kernel for x86, ARM, and RISC-V, and builds QEMU, FFmpeg, SQLite, PostgreSQL, and Redis. It also compiles Doom, which Carlini used as his own litmus test.

The limitations are listed with the same specificity. There is no 16-bit x86 code generator, so booting in real mode is handed to GCC. The integrated assembler and linker were attempted and are still buggy, so those are external too. Generated code is slower than GCC with optimisations turned off. Not every project compiles reliably. Nobody should read this as a drop-in replacement, and the author does not claim it is.

How sixteen agents shared one repository

The coordination scheme is almost embarrassingly simple, which we think is the point. Each agent runs in its own Docker container against a shared git repository. Task locking is a text file. Merge conflicts get resolved by the agents when they arise. The scaffold is a loop that resumes work whenever an agent finishes or stalls, with no orchestrator deciding who does what.

Parallelism paid off when work could be sliced into small independent pieces, such as individual failing tests from a compiler test suite. It struggled on monolithic goals like getting the whole kernel to compile, where sixteen agents would step on each other or chase the same failure. The way out was differential testing against GCC, which turns one large goal into many small discrepancies, each of which is a task an agent can lock and fix.

Two design choices addressed properties of the model rather than the problem. Context pollution, where a long session fills with stale details, was managed by keeping sessions short and restarting. What Carlini calls time blindness, the model's poor sense of how long it has been working or how much remains, was handled by sampling progress and by deterministic subsampling of the test suite so parallel agents did not all run everything.

The tests decide what gets built

The line from the post we would put on a wall is the advice to write extremely high-quality tests, because autonomous agents solve whatever problem the tests define. The compiler passes about 99 percent of most suites it was run against, including the GCC torture tests. That number is a fact about the tests as much as about the compiler. A hole in the suite is a hole in the product, and with agents there is no engineer's intuition filling the gap.

This is the part of the story that transfers beyond compilers. If the work is going to be done by a loop that runs until the checks pass, then writing the checks is the engineering. The agents' contribution was to make the cost of the loop low enough that it is worth being careful about the checks.

The other post from the same day

Gian Segato and colleagues published a second piece on February 5 about infrastructure noise in agentic coding evaluations. They ran the same model with the same scaffold on the same Terminal-Bench 2.0 tasks under six resource configurations, from strictly enforced one-times allocations to uncapped. The spread between the most and least resourced setup was six percentage points, significant at p below 0.01. Infrastructure error rates went from 5.8 percent under strict limits to 0.5 percent uncapped.

A crossover on SWE-bench, 227 problems with ten samples each at five times the baseline RAM, moved scores by 1.54 points. Their recommendation is to specify both a guaranteed allocation and a hard kill threshold per task rather than one pinned number, to calibrate the headroom so that floor and ceiling scores fall within noise, and to run evaluations across different times and days. Their headline is that leaderboard differences below three points deserve scepticism until the configuration is documented. They note the trends appear to hold for models other than Claude but have not tested that rigorously.

Why we read these together

Six points on Terminal-Bench is more than the gap between many adjacent entries on that leaderboard. And the compiler result was produced by a scaffold with locking, restarts, subsampling, and differential testing, none of which is a property of Opus 4.6. Put the two side by side and the lesson is that for long-running work, the environment and the scaffolding around the model are now a larger and cheaper lever than which model you call.

Models still matter. The claim is about where the marginal effort goes. If we had a fixed budget to improve a long-running agent this quarter, we would spend it on container provisioning, test quality, and the loop, and we would expect a bigger return than from a model upgrade. We would also want to see the compiler experiment rerun with a different frontier model on the same scaffold, because that is the only way to find out how much of the 100,000 lines the scaffold deserves credit for.

Sources

  1. Nicholas Carlini, Building a C compiler with a team of parallel Claudes (Anthropic Engineering, February 5, 2026)
  2. Gian Segato et al., Infrastructure noise in agentic coding evals (Anthropic Engineering, February 5, 2026)