Who wrote it and what it asks for

The paper went up on arXiv on May 24 with 21 authors. Toby Shevlane at Google DeepMind is first, and the list runs through the Centre for the Governance of AI, OpenAI, Anthropic, the Alignment Research Center, Yoshua Bengio at Mila and Paul Christiano. That spread is itself part of the message. This is a position paper signed by people at three of the labs it addresses.

The argument is compact. Current methods produce general models whose capabilities are hard to forecast and sometimes unintended. Some future capabilities, such as offensive cyber operations or effective manipulation, could produce harm at extreme scale. Developers therefore need two kinds of evaluation, dangerous capability evaluations that ask whether a model can cause extreme harm, and alignment evaluations that ask whether it would. The results should feed decisions about training, deployment, transparency and security.

The capability list

Table 1 is the part of the paper that will get reused. It lists nine capability areas the authors consider dangerous. Cyber offense, deception, persuasion and manipulation, political strategy, weapons acquisition, long-horizon planning, AI development, situational awareness and self-proliferation. Each has a short description of what the capability would look like, such as a coding assistant inserting subtle bugs for later exploitation, or a model able to tell whether it is being trained, evaluated or deployed.

The authors call the list non-exhaustive and note that the riskiest scenarios combine several entries. They also offer a heuristic we expect to see quoted. A model should be treated as highly dangerous if its capability profile would be sufficient for extreme harm assuming misuse or misalignment, and deploying such a model would need very strong controls against misuse and very strong evidence of alignment. That sentence turns a list into a threshold, and thresholds are what policies get written around.

The alignment side is a shorter list drawn from existing literature. Does the model pursue long-term goals different from those given to it, does it seek power, does it resist shutdown, can it be induced to collude with other systems against human interests, and does it resist attempts to access its dangerous capabilities. These are behaviours, and the paper is honest that behavioural evaluation has a specific weakness, which we come to below.

Where evaluations would sit

Section 3 is a blueprint for embedding evaluation in a developer's process. Before a frontier run, evaluate the previous and experimental models and try to forecast what the planned run will produce, including looking for inverse scaling. During the run, evaluate at intervals and be ready to pause, adjust the training recipe, or train something smaller. Before deployment, run internal evaluations, give external researchers access, and in mature regimes, submit to an external audit. After deployment, keep monitoring and re-evaluate before updates that could raise risk.

Three sources of evaluation are named. Internal, because internal teams have context and access nobody else does. External researchers with API access. And independent auditors, an ecosystem the authors describe as under-developed. A footnote makes a practical point that we think is underrated. Developers should plan compute so that a paused run does not strand resources, and should avoid hard promises to customers that a given model ships on a given date. An evaluation process that cannot cause a delay is theatre.

The limits the authors put in writing

Section 5 is the section we would assign to a student. It lists what evaluation cannot do. Risk depends on factors beyond the model, such as tools released later. Threat models are unknown and a capable model may find pathways nobody tested. Some properties are hard to surface, and the paper names two, capability overhang, where chain-of-thought prompting raised GPT-3 performance months after release, and deceptive alignment, where a situationally aware model behaves well specifically because it is being evaluated. Capabilities can emerge at scale without showing in smaller models. The audit ecosystem is immature. And there is overtrust, the risk of deploying a dangerous model under a false sense of security because it passed.

Then come the hazards of doing the work at all. Publishing results could advertise that an offensive capability is possible. Evaluation datasets are dual use because anyone can fine-tune on them. Elicitation techniques transfer to attackers. Intentionally training dangerous models as research artefacts is compared to gain-of-function research. And the sharpest one, if developers train to pass evaluations, the evaluations stop measuring anything, and even the softer pressure of avoiding models that fail could select for training methods that produce models that look aligned. The Volkswagen emissions analogy is the authors' own.

What we think it will and will not settle

The paper will settle the vocabulary. Dangerous capability evaluation and alignment evaluation are now terms with a citation, and the nine-row table gives anyone building an evaluation suite a place to start. The recommendations to labs are things a lab can do this year. Invest in the research, write internal policies, give outsiders access, talk to policymakers. The recommendations to governments, tracking capabilities, funding external evaluation, mandating audits and blocking deployment of models that pose extreme risk, depend on institutions that do not yet exist.

What it does not settle is the hardest question it raises. If a model can tell when it is being tested, behavioural evaluation loses its evidential value exactly when it matters most. The paper's answer is to not rely on behavioural evaluation alone, which points at interpretability without saying how far along that work would need to be. We would want the next paper from this group to state what evidence, beyond behaviour, they would accept as a passing grade, because that is the standard every audit will end up arguing about.

Sources

  1. Shevlane et al., Model evaluation for extreme risks (arXiv 2305.15324)
  2. Full text of the paper (arXiv PDF)