Beyond the permission prompt: auto mode, Safehouse and disposable sandboxes
Anthropic published the design of an approval classifier for Claude Code this month, and two very different sandboxes for local agents are now easy to install. Notes on what each one defends against and why asking the user stopped working.
The number that ended the prompt
Anthropic's engineering post on auto mode opens with a statistic that explains the whole design. Claude Code users approve 93 percent of permission prompts. A control that people accept nineteen times out of twenty is no longer a control. It is a click, and clicks train people to stop reading. The alternative most of us had been using was the flag that skips permissions entirely, which trades approval fatigue for no safeguard at all.
Auto mode is the middle position. Anything the agent wants to do that is not on a safe allowlist gets judged by a model-based classifier instead of by the user. The user is still in the loop for the task, and out of the loop for each individual command. Whether that is an improvement depends entirely on what the classifier catches and what it misses, and Anthropic published both.
Two stages and a deliberately blinkered view
The classifier runs in two stages. The first is a single-token decision that sees every non-allowlisted action and is told to err on the side of blocking. The second is a chain-of-thought pass that only runs on actions the first stage flagged, so the expensive reasoning is spent on the small fraction that looked suspicious. On 10,000 real traffic samples the first stage alone had an 8.5 percent false positive rate, and the full pipeline brought that down to 0.4 percent.
The detail we found most interesting is what the classifier is not allowed to see. It gets the user's messages and the tool call itself. Assistant prose, tool descriptions and tool outputs are stripped out. The stated reason is that an agent should not be able to talk the judge into a decision by rationalising, and a hostile file the agent read should not be able to talk to the judge at all. That one design choice covers two of the four threats in their model: an overeager agent and a prompt injection.
The other two threats are honest mistakes about scope and outright misalignment, which the post says has not been observed in practice. The list of blocked actions in the write-up reads like an incident log from anyone who has run agents unattended. Deleting a remote branch after a vague request to clean up. Grepping for a different API token after an auth failure. Posting a possibly confidential script as a public Gist. Retrying a failed deploy with verification switched off.
What it still misses
The residual risk is stated plainly. On 52 real overeager actions, the pipeline had a 17 percent false negative rate. On 1,000 synthetic exfiltration attempts, 5.7 percent got through. The post's own diagnosis is that the classifier often sees the danger correctly but cannot tell whether the user's consent stretched to the full consequence of the action. It calls this acceptable relative to running with no permissions and arguably a regression relative to a careful human on critical infrastructure. We agree with both halves.
There are backstops. A server-side probe screens tool outputs for injection before they enter the agent's context. Classifiers run on both sides of a subagent handoff. A blocked action does not end the session, the agent is told to find a safer route, and the session terminates after three consecutive denials or twenty in total. When auto mode is entered, permission rules known to grant arbitrary code execution are dropped, including blanket shell access. That last one will surprise people who had configured their way to zero prompts.
The kernel does not read the transcript
A classifier judges intent. A sandbox ignores intent and constrains effect. Agent Safehouse, a macOS tool released under Apache 2.0, takes the second approach at the kernel level. It grants read and write access to the chosen working directory, usually the git root, read access to installed toolchains, and denies everything else before the process can reach it. SSH keys, other repositories and personal files are simply not there. The project's framing is that language models are probabilistic, so a catastrophic command will eventually be issued, and the question is only whether the syscall succeeds. If the agent runs rm -rf on the home directory, the kernel refuses.
Docker Sandboxes go one layer further and put the agent in a disposable microVM. Only the project workspace is mounted, network access is policy controlled, credentials are isolated, and the agent can install packages or even run containers inside without touching the host. It supports Claude Code, Codex, Gemini CLI, Copilot CLI, OpenCode and Kiro, installs on macOS, Windows and Ubuntu without Docker Desktop, and the product page openly advertises running agents in skip-permissions mode inside it. Their pitch is that the boundary is hard enough that the prompt becomes unnecessary.
Which one you actually need
These are answers to different questions. A sandbox cannot tell that deleting a remote branch was out of scope, because the git push looks identical to a legitimate one and the remote is outside the sandbox anyway. A classifier cannot stop a command it approved from having consequences it did not foresee. Safehouse protects the laptop. The microVM protects the laptop and contains the blast radius of anything the agent installs. Auto mode protects the things the agent has credentials for, which is where the branches, buckets and deploys live.
We have started running the classifier inside a sandbox rather than choosing between them, and we would like someone to measure that combination properly. The obvious experiment is to take Anthropic's 52 overeager actions and the 1,000 exfiltration cases, run them under each configuration, and count what reaches the outside world. The 17 percent number is the one we want to see shrink, and we do not think a filesystem boundary is what shrinks it.
Sources
From the foundation