RSP v3: the scaling policy admits it cannot go it alone
Anthropic's third Responsible Scaling Policy separates what the company will do on its own from what it thinks the industry should do, replaces hard commitments at the higher levels with a graded roadmap, and adds Risk Reports with external review. Something was gained and something was given up.
What the company says went wrong with v2
The announcement is candid about three problems. First, the capability thresholds turned out to be far more ambiguous than anticipated. On biological risk in particular, models now pass most of the tests, but passing a test does not settle whether a model provides real uplift, and the wet-lab trials that could settle it take long enough that a more capable model ships before they finish. Second, the policy was written on the assumption that governments would respond as capabilities rose, and the company now says policy has moved toward competitiveness and growth instead. Third, the mitigations required at ASL-4 and above may be impossible for one company to implement alone. The example given is weight security against state-level attackers, which a RAND assessment described as currently not possible without help from the national security community.
The original RSP from September 2023 was built as an if-then document. If a model crosses a threshold, then stricter safeguards apply before it can be deployed or trained further. The implied consequence was a pause. Version 3 keeps the if-then shape for what the company can actually do and drops it for what it cannot.
The three-part structure
The new document separates into three pieces. The first is a set of unilateral commitments that Anthropic says it will keep regardless of what others do. The second is an industry-wide map from capabilities to mitigations, which is a recommendation rather than a promise, describing what the company thinks every frontier developer should do at each level. The third is a Frontier Safety Roadmap, a list of nonbinding but public targets that the company will grade itself against in the open. Examples in the announcement include an R&D programme on information security beyond current practice, red-teaming methods that go beyond a bug bounty with hundreds of participants, measures to make Claude follow its constitution, centralised records of critical AI development activities, and a policy roadmap with a proposed regulatory ladder.
Alongside that, the company will publish Risk Reports every three to six months that cover capabilities, threat models and mitigations together with an overall risk assessment, with minimal redactions for IP, legal, safety and privacy reasons. External experts with no conflicts of interest are to review those reports with unredacted or minimally redacted access when warranted. The review is being piloted and is not yet required for existing models.
What was gained
The Risk Reports are the concrete win. A document every few months that states what the model can do and what has been done about it, with an outside reader who has seen the unredacted version, is more than any lab has published to date on a schedule. If the reports are as detailed as promised, they give people outside the company something to argue with. The roadmap, graded publicly, is a weaker form of the same thing. It turns a promise into a scoreboard.
Holden Karnofsky, who joined Anthropic full time in January 2025 and says he began pushing for this revision about a year ago, argues for the change on grounds we find partly convincing. A commitment you cannot articulate well may push a company toward the wrong action or toward quiet noncompliance, and a public risk assessment with no automatic trigger attached can be written honestly, because the author has no incentive to shade the finding. He also writes that in the current political environment a unilateral pause would simply make the company look deluded and over-alarmist. That is a claim about optics, and we note that it is being offered as a reason to change a safety policy.
What was loosened
The language that suggested Anthropic would unilaterally stop if it could not meet its own standards is gone. So are the binding commitments to reach state-level security on a compressed timeline, and the rigid requirements for ASL-4 and ASL-5 preparation that leadership judged infeasible within the two years in which it expects to hit the CBRN-4 and AI R&D-5 thresholds. The company would say these were never realistic and that removing them is honesty. The critics on the forum thread would say they were promises.
The sharpest version of the objection comes from a commenter who reports that Anthropic employees told him on more than a dozen occasions that the RSP bound them to a mast. If the policy was sold internally and externally as a binding constraint, and the constraint is removed at the moment it would start to bind, then the escape clause was doing more work than anyone admitted. Owen Cotton-Barratt, on the same thread, is relieved, on the grounds that a policy that cannot be kept is worse than one that says what it means. We think both readings are correct at once, which is the uncomfortable part.
What we would want next
The structure of v3 makes the external review the load-bearing element. A self-graded roadmap and a self-written Risk Report are only as good as the outsider who reads them. So the details we want are boring ones. Who the reviewers are, what they are allowed to publish if they disagree, and whether a disagreement has any consequence beyond a footnote.
The industry-wide map is the other thing to watch. It is an invitation to other labs and to regulators to adopt the same capability-to-mitigation table. If nobody takes it up, the map is a description of what one company thinks everyone should do while doing less itself. If it becomes the basis for the frameworks that California now requires developers to publish, then the concession in v3 will have bought something. Either way, we will know by the time the second Risk Report comes out.
Sources
From the foundation