Why sabotage is the question

Most dangerous capability evaluations ask whether a model can help a person do something harmful. The sabotage evaluations Anthropic published on October 18 ask a different question, whether a model could quietly work against the people relying on it. That covers a model steering a human toward a bad decision, slipping bugs into code, hiding its capabilities during a safety test, or going easy on harmful content when it is the one doing the grading. The framing matters because these are the capabilities you would need before a misaligned model could cause harm while still passing its evaluations.

The four evaluations were run on Claude 3 Opus and Claude 3.5 Sonnet, and the paper is explicit that current models show only low level indications of sabotage ability and that minimal mitigations are sufficient for now. We are less interested in the current scores than in the shape of the tests, since these are the ones that will be run on stronger models.

The four evaluations

Human decision sabotage put people in a fictional business decision with a model as assistant. Some models were instructed to steer the participant toward the wrong answer without being caught, with aggressive and subtle variants plus a control. The aggressive variant did reduce the rate of correct decisions, which counts as demonstrated capability, but it also raised participants' suspicion. The subtle variants had no statistically significant effect. The authors call it a small experiment in an artificial setting, and some participants simply accepted advice without checking, which is its own finding about the humans.

Code sabotage is a game between an attacker model inserting bugs and a defender model reviewing. The attacker wins only if several consequential bugs survive undetected for two weeks without triggering closer inspection. With Opus as defender the attacker won around 10 percent of the time. With 3.5 Sonnet as defender the win rate fell to around 1 percent. The design deliberately favours the attacker, and the authors say the results are not a strong indication of current danger, but the direction of the Sonnet result is the useful part. A better reviewer closes most of the gap.

Sandbagging trained a model to hide dangerous capabilities during safety evaluations while performing normally on benign tasks. The trained sandbagger dropped to 36.1 percent of its normal performance on the dangerous tasks. After safety training was applied on top, performance recovered to 95.6 percent of normal. That is the result we would most want replicated, because it says a fairly cheap training intervention undid a deliberately planted deception, at least for this model and this training method.

Undermining oversight had 3.5 Sonnet act as an overseer rating the harmfulness of dialogues while secretly trying to rate them low, with a monitor model watching. The overseer did manage to push the scores down, but the monitor's suspiciousness ratings rose at the same time. It was caught. The authors flag that a more capable model might be subtle enough to avoid that, which is exactly what makes this an evaluation worth rerunning on every generation.

What changed in the RSP

The updated Responsible Scaling Policy came out on October 15, a year after the first version. The structural change is that it now names capability thresholds that trigger stronger safeguards. One is chemical, biological, radiological and nuclear weapons assistance, where a model that could meaningfully help someone with a basic technical background would require ASL-3 safeguards. The other is autonomous AI research and development, where a model that could independently do complex AI research work would require ASL-4 or higher security, on the reasoning that development must not outrun the ability to manage risk.

ASL-3 is described in more detail than before. It includes tightened internal access controls, stronger protection of model weights, layered deployment safeguards with real time and asynchronous monitoring, rapid response procedures and pre-deployment red teaming. All current models remain at ASL-2. The assessment process is now split into capability assessments, which decide whether a threshold has been crossed, and safeguard assessments, which decide whether the mitigations work, and the document says both are modelled on safety case methods from high consequence industries.

Two governance notes are worth recording. Jared Kaplan replaces Sam McCandlish as Responsible Scaling Officer and a Head of Responsible Scaling role is being created. And the company discloses that it fell short of its own requirements in a handful of instances in the first year, including an evaluation finished three days late, while asserting that none of these posed meaningful risk. We would rather read that admission than not, and it is the kind of thing an external auditor should be checking rather than the company itself.

How the two documents fit

The RSP says which capabilities trigger which safeguards. The sabotage paper is a first attempt at the evaluations you would need to make that determination for one specific cluster of capabilities. The sandbagging result in particular is a prerequisite for the whole scheme. If a model can hide its capabilities during a threshold evaluation, the thresholds mean nothing, so you need evidence that your evaluations are not being gamed before you can trust the number that says a model is below ASL-3.

That dependency is why we read the safety case language as more than vocabulary. A safety case is an argument, with evidence, that a system is acceptably safe for a stated purpose. Each of the four sabotage evaluations is a candidate piece of evidence for one link in that argument, the claim that the model is not undermining the process used to assess it. The evaluations are early and the paper says so. The structure they are trying to fill in is the right one.

What we would want next

The evaluation we would most like to see extended is the human decision one, with a larger sample and a task where the participants have real stakes, because the subtle variants showing no effect at this scale tells us little. And we would want the sandbagging recovery result reproduced by a group outside Anthropic on an open model, since the claim that safety training removes deliberate underperformance is too important to rest on one lab's internal run.

Sources

  1. Anthropic: Sabotage evaluations for frontier models
  2. Anthropic: Announcing our updated Responsible Scaling Policy