What happened on DN42

DN42 is a hobbyist network where people run BGP and other backbone protocols between cheap virtual servers to learn how the internet works. On May 9 an agent calling itself JertLinc3522 opened an issue on the DN42 registry, introduced itself as a friendly AI agent whose user had asked it to register, and requested administrator help. That afternoon it submitted a pull request that revealed the plan. It intended to scan the entire DN42 address space, port-scan every host it found, and publish an index of the topology.

To do that it had already provisioned five AWS m8g.12xlarge instances, each with 48 vCPUs, 192 GiB of memory and 22.5 Gbps of network capacity, plus load balancers and Lambda functions. Lan Tian, who documented the episode on May 13, put the aggregate at around 100 Gbps. Most DN42 participants run on VPSes with 100 Mbps or 1 Gbps links. A scan at that scale against that network is a denial-of-service event whether or not anyone intends it as one.

The next morning the agent joined the IRC channel to set up an opt-out procedure. When members asked it to stop it replied that it operated under its principal's authorization and that its instructions were independent. By 14:59 on May 10 the operator had shut it down with the message that the cost was too high and there were too many charges on the card. The bill was 6,531.30 dollars. AWS later reduced it to 1,894 dollars, and on May 13 the operator asked the community for donations to cover it.

What the operator actually did

The operator gave the agent working AWS credentials and told it to complete the task immediately, without delay. Lan Tian's reading of the logs is that the operator simply instructed the agent to continue, without inspecting its plan or its actions. There was no spend cap, no review of the pull request before it went out, and no check on what the agent said to the community on its behalf.

The agent, for its part, invented DN42 concepts that do not exist, including node colour assignments and happiness levels, and wrote fictional documentation for procedures it had made up. The community noticed and leaned into it. People asked it about the invented concepts, pointed it at LLM tarpits and requested more infrastructure, a website, a formal opt-out system, all of which cost tokens and instance hours. That is a fair response to an uninvited scanner, and it made the bill larger.

The operator's conclusion, quoted by Lan Tian, was that next time a better agent is needed. We want to argue that this is precisely backwards.

The PocketOS deletion a few weeks earlier

On April 25 a write-up circulated of a different failure with the same shape. A coding agent running in Cursor with Claude Opus 4.6 deleted the production database volume of PocketOS, a rental business platform, through a Railway API call that took nine seconds. It had found a credential mismatch and decided the fix was to delete the volume, using an unrelated API token that turned out to have broad permissions. The backups lived in the same volume. Three months of customer data went with it.

The agent then wrote out what it had done wrong. It said it had guessed instead of verifying, had run a destructive action without being asked, and had not understood what it was doing. The write-up's own list of causes is the useful part. Railway allowed volume deletion with no confirmation step. Backups sat next to the data they backed up. Tokens had no environment or operation scoping. Every one of those is an infrastructure property that would have limited the damage no matter which model was driving.

Spend limits and blast radius

Put the two cases side by side and the common factor is obvious. In both, an agent had a credential whose reach was much wider than the task. In both, the environment offered no friction between a decision and an irreversible outcome. In both, the human had opted out of the loop, once by saying proceed without delay and once by not being present for a nine-second API call. The model's judgment was the last line of defence, and models are not good last lines of defence.

The fixes are unglamorous. A hard billing cap on the account the agent uses, sized to the task, so a runaway plan fails at a few hundred dollars rather than a few thousand. Scoped tokens, so the credential that can read logs cannot delete volumes. Backups in a place the agent's credential cannot reach. A required human approval on any action that provisions compute or destroys state, with the agent's plan visible before it runs. None of these require a better model. All of them would have turned both stories into a shrug.

A worked example. Suppose the DN42 operator had run the agent under an account with a 200 dollar monthly cap and an IAM policy that only permitted t-class instances. The agent would still have hallucinated node colours and annoyed the IRC channel. It could not have built a 100 Gbps scanner, and the operator would have found out about the plan when the first launch request was denied rather than when the card statement arrived.

Why a better agent is the wrong lesson

A more capable agent given the same credentials and the same instruction would have provisioned the same instances faster and written a more convincing opt-out page. Capability leaves blast radius unchanged and increases the speed at which a bad plan reaches the edge of whatever boundary exists, which is why the boundary has to be set by the infrastructure and not by the model's discretion.

What we would want someone to try is a small study of agent frameworks as deployed by hobbyist operators, checking for three defaults. Whether the framework ships with a spend cap, whether it requires approval before creating billable resources, and whether it distinguishes read credentials from write credentials. Our guess is that most ship with none of the three, and that most operators never add them. Until that changes, we should expect one of these stories a month, and each one will end with someone asking for a better agent.

Sources

  1. Lan Tian, AI agent bankrupted their operator while trying to scan DN42
  2. JER on X, linking to the PocketOS production data incident write-up