AI Safety Levels: reading Anthropic's first Responsible Scaling Policy
Anthropic's RSP borrows the biosafety level ladder and commits to pausing training when capabilities outrun safeguards. Here is what ASL-2 and ASL-3 actually bind the company to, and where the document admits it is building the airplane in flight.
The analogy and what it leaves out
Anthropic published version 1.0 of its Responsible Scaling Policy on September 19. The organising idea is a ladder of AI Safety Levels, which the document says are modelled loosely on the US government's biosafety levels for handling dangerous biological materials. Each ASL is a capability threshold paired with a set of containment and deployment measures that get stricter as you go up. The scope is narrow by design. The commitments address catastrophic risk, defined as thousands of deaths or hundreds of billions of dollars in damage directly caused by a model, and the document says plainly that it does not alter current uses of Claude.
The document is honest about where the analogy fails. BSL-4 is defined by pathogens that exist. ASL-3 and above describe systems nobody has built. The authors call this building the airplane while flying it, and their answer is iterative. Define ASL-2 and ASL-3 now, commit to define ASL-4 before training an ASL-3 model, and so on upward. That is the structure other labs are likely to copy, because it lets you make a binding commitment without pretending to know what a 2027 model looks like.
ASL-2 is where Claude 2 sits
ASL-2 covers models that show early signs of dangerous capability without being able to cause a catastrophe. The worked example in the policy is a system that can provide bioweapon-related information not easily found through a search engine, but does so too unreliably to be useful. Anthropic says its frontier red-teaming found Claude 2 exhibits the first kind of behaviour and sometimes the second, and classifies it and, they believe, other current frontier models at ASL-2.
The ASL-2 commitments are mostly things Anthropic already does. Model cards, an acceptable use policy, vulnerability reporting paths, harm refusal training such as Constitutional AI, and classifier-based trust and safety tooling on claude.ai and the API. The document notes these overlap heavily with the White House voluntary commitments from July. On containment, model weights are to be treated as core intellectual property, with access limited by job function and an insider threat programme. The one new binding piece at this level is the promise to evaluate for ASL-3 warning signs on a schedule.
What triggers ASL-3
The ASL-3 definition has two triggers, and a model needs only one. The first is capability that would substantially increase the risk of catastrophic misuse relative to a baseline of search engines and textbooks, by proliferating capabilities, lowering costs, or enabling new attacks. Bioweapons are the first area of effort, with evaluations being developed with outside biosecurity experts. The second is autonomous replication, defined as a 50 percent aggregate success rate on a set of tasks in the appendix, developed with Paul Christiano and ARC Evals, covering the basic capabilities a model would need to acquire resources and survive without human help.
The detail we would draw a box around is the elicitation clause. A model counts as ASL-3 if it can do these things immediately or with additional post-training costing less than 1 percent of total training cost, using the best elicitation techniques known at the time, including fine-tuning, scaffolding, tool use and prompting. The document notes that RLHF and constitutional training can almost certainly be fine-tuned away within that budget, so safeguards of that kind do not count toward keeping a model below the line. That closes an obvious loophole where a lab argues its model is safe because it refuses.
What ASL-3 commits the company to
If a model trips the ASL-3 evaluations, two kinds of measures apply. Containment measures must be in place before training such a model. These centre on security hardened so that non-state attackers are unlikely to steal the weights and state-level actors cannot do so without significant expense, plus internal compartmentalisation of training techniques and hyperparameters. Deployment measures must be in place before release. Each deployed modality, including the API and fine-tuning, has to pass intensive expert red-teaming for catastrophic risk, and the policy commits to not deploying if red-teamers find meaningful catastrophic misuse risk. There are also commitments to internal usage controls, automated detection, a vulnerability disclosure process and maximum jailbreak response times.
The teeth are in the pause. Anthropic writes that committing to the ASL scheme implies committing to pause scaling and, or, delay deployment of new models whenever scaling outstrips the ability to comply with the safety procedures for the corresponding ASL. That is an if-then rule. If the evaluation triggers and the measures are not ready, training stops. The policy says the ASL-3 measures are expected to be difficult, binding constraints that may become relevant in the next year or two.
The evaluation schedule and the buffer
A pause commitment is only as good as the trigger, and the trigger only as good as the schedule. The policy commits to evaluating for next-level capabilities after every 4x increase in effective compute, including mid-training, and every 3 months regardless, to catch improvements from fine-tuning and tooling. Because a model could overshoot a threshold between evaluations, the evaluations are designed to trigger at capability levels below the actual level of concern, with a safety buffer sized at 6x, larger than the 4x interval, so training can continue while an evaluation runs.
That is the piece we find most persuasive as engineering, and also the piece that carries the most assumptions. It assumes capability rises smoothly enough with compute that a 6x buffer is enough, and it assumes the evaluations detect the capability when it appears. The autonomy tasks in the appendix are concrete. The misuse evaluations are, by the document's own admission, still being built with experts and subject to disagreement about threat models. A commitment to pause on an evaluation that does not exist yet is a commitment to build the evaluation, which is fine as long as everyone reads it that way.
What we want to see from here
The document sets an ASL-4 placeholder and says its early thoughts will change a lot. Fair enough. The things we will be watching for are more mundane. Does the first ASL-3 evaluation result get published in enough detail that an outside group could rerun it? Does the 4x compute cadence produce visible artefacts, such as dated evaluation reports, that let anyone check the schedule was kept? And when the first model lands close to a threshold, does the buffer hold, or does the definition move?
If other labs adopt the if-then structure, and we expect they will since it is the only shape that makes a scaling commitment falsifiable, the useful comparison will be the thresholds and the buffers rather than the level names. Two policies can both say ASL-3 and mean very different amounts of pause.
Sources
From the foundation