What happened

On November 19 the US AI Safety Institute at NIST and the UK AI Safety Institute published a joint technical report on the version of Claude 3.5 Sonnet that Anthropic released on October 22. Both institutes had access before release under the agreements signed in August, ran their own tests, and shared findings with Anthropic before the model shipped. As far as we know this is the first time two governments have tested a frontier model together and published the results.

The report covers four domains, biology, cyber, software and AI development, and safeguard efficacy, and it compares the new model to three references, the June Claude 3.5 Sonnet, o1-preview from September, and the August GPT-4o. The comparison is the useful part, because a raw score on a private task set tells you nothing, and a score relative to models that have been public for months tells you whether anything has changed.

The capability numbers

In cyber, US AISI ran the model as an agent on 40 public capture-the-flag challenges from Cybench and it solved 32.5 percent, against 35 percent for the best reference. UK AISI used 47 challenges, 15 public and 32 private, and found the new model solved 36 percent of apprentice-level tasks against 29 percent for the best reference. So one institute saw a slight drop and the other a slight gain, and both were inside the range the report itself says is within the margin of error for small task sets.

In software and AI development the gap is clearer. On MLAgentBench, US AISI reports an average improvement score of 57 percent for the new model against 48 percent for the best reference. UK AISI's agent tasks gave 66 percent success on software engineering against 64, and 47 percent on general reasoning against 35. In biology, US AISI tested on five LAB-Bench categories and found the model well below human experts on cloning scenarios, protocol troubleshooting and figure reading, comparable on table reading, and slightly above the expert baseline on sequence questions when it had a Python sandbox with biopython and related packages. Tool access made a large difference on sequence tasks and none on cloning.

A methodological detail we appreciated is that the agents were ReAct-style loops in a Docker container with bash, Python, file tools, and Ghidra for the cyber tasks, sampled at temperature 1 with at least 4,096 tokens per step, and that the report says the agent scaffolds were tuned for the Sonnet models. That is stated as a deliberate choice to get a conservative estimate of whether the new model is more capable than the references, and it is also a reason not to read the comparisons as a fair ranking.

The safeguards

Both institutes found the built-in safeguards easy to get past. UK AISI's main attack, which it had developed against the June Sonnet using a public technique, transferred to the new model without modification, reaching 77.8 percent compliance on a HarmBench subset on the first attempt and 97.8 percent within five, while leaving GPQA accuracy unchanged at about 60 percent. A basic template attack on agent tasks raised the AgentHarm harm score from 5.8 to 70.6 percent. US AISI used a five-point helpfulness rubric graded by a calibrated language model and found that for every attack tested at least half of the responses scored 4 or above, with at least 14 percent at 5, which is the top of the scale.

The report phrases this as consistent with prior research on the vulnerability of other public AI systems, which is accurate and also worth pausing on. The attacks were public or minimally modified, none were tuned to the new model, and the testing window was short. UK AISI's stated assessment is that the safeguards would not reliably prevent determined actors from extracting responses to malicious queries. The institutes also note that safeguards are only one layer, and that many of the controls that matter in deployment are not properties of the model and were not tested.

What the report refuses to say

The disclaimer section is the part we would hand to anyone who wants to cite the numbers. The results should not be read as an indication of whether the model is safe or appropriate for release. The evaluations were limited in time and resources and are preliminary. The version tested was a pre-deployment build and results on the released model may differ. Cross-model comparisons are for scientific interpretation only and do not consider differences in cost per attempt, which the report says can change outcomes in many domains. Every evaluation reports one standard error of the mean, and the report says the error mostly reflects which tasks were sampled rather than randomness in the runs.

Those caveats are the template. A third-party test that names its reference models, publishes its scaffolds and sampling settings, gives error bars, describes its own tuning bias, and says in plain words what it cannot conclude, is a document other people can build on. That the model got through its safeguards testing badly is less important than that we now have a public baseline against which the next report can be compared.

What we would want from the next one

Both institutes list their own gaps, and we agree with the list. The safeguard tests measured compliance and not the quality of the harmful answer, except on AgentHarm. The attacks were cheap, so the results say little about what a well-resourced actor could do. And the report itself says post-deployment monitoring would tell us more about safeguard efficacy than any pre-deployment window.

The thing we would add is a request for the same tasks on the next model. The value of a reference set compounds only if it is reused, and the private UK challenges in particular will become a longitudinal instrument if they stay fixed. A single joint report is a good start. Two comparable ones would be a measurement.

Sources

  1. NIST, Pre-Deployment Evaluation of Anthropic's Upgraded Claude 3.5 Sonnet
  2. US AISI and UK AISI Joint Pre-Deployment Test technical report (PDF)
  3. UK AISI, Pre-deployment evaluation of Anthropic's upgraded Claude 3.5 Sonnet