ASL-3 turned on: what a provisional safety level actually commits a lab to
Anthropic activated its ASL-3 deployment and security standard for Claude Opus 4 without concluding the model needs it. A reading of the report on what triggered the decision, what the measures cover, and what provisional means in practice.
A trigger that was not a threshold
On May 22 Anthropic released Claude Opus 4 under the ASL-3 deployment and security standards from its Responsible Scaling Policy, the first model it has shipped that way. The interesting sentence in the announcement is that the company has not determined the model requires ASL-3. It activated the standard because it could not rule out the need for it.
The report explains the reasoning. The relevant capability threshold, CBRN-3, is defined as the ability to significantly assist people with basic STEM backgrounds in obtaining, producing or deploying chemical, biological, radiological or nuclear weapons. Claude Sonnet 3.7 was far enough below that line to ship under ASL-2, but showed enough improvement on biological weapons proxy tasks that the company decided its next flagship would get the heavier protections regardless. Opus 4 then performed clearly better than Sonnet 3.7 on some of those proxy tasks, and external red teamers reported it behaved qualitatively differently from anything they had tested. The report also notes that dangerous capability evaluations take longer to settle as a model approaches the line, which is why the decision is described as precautionary and provisional. Anthropic says it may return the model to ASL-2 if evaluation concludes the baseline protections would have sufficed. Claude Sonnet 4, released the same day, showed more modest gains and stays under ASL-2.
What the deployment standard covers
The deployment side is narrower than the label suggests. It is aimed at universal jailbreaks, meaning prompting strategies that reliably get past both the model and its surrounding classifiers, and specifically at jailbreaks that would let someone extract help with a long, multi-step weapons workflow. The report is explicit that the measures are not meant to block single pieces of commonly available information, and gives the chemical formula for sarin as an example of a question they are not designed to stop, though they often do.
The main control is a set of real-time classifiers built on the Constitutional Classifiers work from January. Anthropic wrote a constitution describing three categories of use, prohibited, permissible but relevant to the threat model, and clearly permissible, generated synthetic prompts and completions for each, and fine-tuned language models to block the first category on both input and output. The initial scope is biological weapons only, with other CBRN categories under evaluation. The report concedes false positives are inevitable, says they have been reduced substantially between research and production, and says the classifiers add moderate compute overhead.
Around the classifiers sit three further pieces. Offline monitoring runs slower, more capable checks on traffic to estimate how often the real-time guards are bypassed and to measure the false positive rate. A bug bounty run with HackerOne, which started as a pre-launch test on Sonnet 3.7, now pays for universal jailbreaks against the deployed classifiers. And a rapid response process aims to patch discovered jailbreaks within days to weeks, partly by having a model generate many variants of an observed attack and training the classifier on them. There is also a vetted access route for organisations doing legitimate dual-use science, which can receive exemptions from classifier actions with ongoing monitoring.
What the security standard covers
The security side is about model weights, and the threat model is bounded. In scope are hacktivists, criminal groups, terrorist organisations, corporate espionage, basic insiders and undifferentiated attacks from state-sponsored groups. Out of scope are sophisticated insiders, nation-state attackers and APT-level actors, which the RSP reserves for ASL-4. The report says the ASL-3 work added more than 100 new controls, and it highlights four.
Egress bandwidth controls limit how fast data can leave the environments where weights live, on the theory that weights are large and a low enough rate limit makes exfiltration slow enough to detect. Anthropic describes these as preliminary and as a step towards possibly airgapping weights later. Two-party control requires anyone accessing weights to authenticate with a physical key, state a justification and get a second person's approval, with time-limited grants. Endpoint controls use binary allowlisting alongside detection and response software. And change management for repositories that touch weights or security infrastructure now requires signed commits, extra review on some changes and designated ownership, with Claude used to speed re-review of routine changes.
None of those four is novel as a security practice, and the report says so. The novelty is in a lab committing to them publicly as a condition of releasing a model, and tying them to a stated threat actor list that can be checked against what the controls plausibly stop.
What provisional means
The word provisional does real work here, and we want to be precise about what it commits the company to. It commits them to running the ASL-3 measures now, at the cost of compute overhead, false positives and engineering friction, on a model they have not concluded needs them. It commits them to finishing the evaluation and publishing the outcome either way. It does not commit them to keeping the measures on, since the report explicitly reserves the option of turning them off.
That last point is the one we would watch. The honest reading of the report is that the company reached the edge of what its evaluations could resolve and chose the conservative side while it kept measuring. The uncharitable reading is that a standard which can be switched off after launch is a standard with a built-in exit. Which reading is right depends on whether the follow-up evaluation is published with enough detail to argue with, and on whether the bug bounty and offline monitoring numbers are released rather than summarised.
What we would want to see next
Three things would make this easier to assess from outside. The first is the false positive rate on real traffic, since a classifier that blocks legitimate biology questions at any meaningful rate is a cost that should be visible. The second is the bug bounty results, in particular how many universal jailbreaks were found, how long each took to patch, and whether the patched versions held. The third is the eventual ruling on whether Opus 4 crossed CBRN-3, with the proxy task results that decided it.
The report ends by saying the implementation will almost certainly not be perfect and that the company expects to debug it in public. That is the right posture. Whether it is more than a posture will be decided by what gets published over the next six months.
Sources
From the foundation