What the paper is

A policy paper appeared on arXiv on October 26 with a title, Managing extreme AI risks amid rapid progress, and an author list that reads like a conference plenary. Yoshua Bengio, Geoffrey Hinton, Andrew Yao, Dawn Song, Pieter Abbeel, Stuart Russell, Daniel Kahneman, Yuval Noah Harari, Gillian Hadfield, Jeff Clune, David Krueger, Anca Dragan and a dozen more. It is short, it is written for policymakers, and it arrived days before the UK AI safety summit at Bletchley Park.

The argument is that capabilities are advancing quickly, that the systems now being built act with increasing autonomy, and that both technical safety research and governance are behind. The risks named are large scale social harm, malicious use, and an irreversible loss of human control over autonomous systems. The remedy proposed is a combination of technical research and what the authors call proactive, adaptive governance, drawing on how other safety critical industries are regulated.

The demands, stated plainly

The headline number is a third. The paper asks major technology companies and public funders to put at least one third of their AI research and development budgets toward safety and ethical use. It supports that with an estimate that only 1 to 3 percent of AI publications are on safety. Whether or not you accept the specific fraction, the contrast between a third of budgets and a few percent of papers is the gap the authors want closed.

The governance list is long and concrete. Mandatory registration of frontier systems and their datasets across the lifecycle. Monitoring of model development and of supercomputer use. Whistleblower protections and incident reporting inside developers. Governments prepared to license frontier development and to restrict the autonomy of these systems in key societal roles. Legal accountability for developers and owners for harms that could reasonably have been foreseen and prevented. And a requirement that developers grant external auditors on site, white box and fine tuning access from the start of development, which is the item we would expect labs to resist hardest.

Why it is not a consensus

The paper is being described as a consensus statement, and in one sense it is. Getting this group to sign one document is an achievement. In the sense that matters for policy, it is not. The author list is drawn from the part of the field that already believes extreme risk is real and near. The researchers who have argued publicly that these concerns are overstated, and there are prominent ones, are not on it. A document that convinces the convinced is a position paper, and it should be read as one.

It is also not a consensus with the labs. The demands fall almost entirely on developers and governments, and none of the frontier labs is represented among the authors. A third of R&D budgets, white box auditor access from day one, and licensing are each a large ask of a company in a race. The paper does not say what happens when the labs decline, and at the time of writing nothing obliges them to do otherwise.

What is missing from the technical side

As a researcher we wanted more on what the third of the budget would buy. The paper names the problem, that we do not know how to make highly capable systems reliably safe, and lists the risks, but it is thin on which technical directions the authors consider most promising and how progress would be measured. That may be appropriate for a document aimed at ministers. It leaves a gap for the research community, because a funding mandate without a research agenda tends to fund whatever already has a name.

There is also a tension the paper does not resolve. It asks for more research on frontier systems and for tighter control over who may build them. Those pull in opposite directions for anyone outside the labs. If the systems worth studying can only be built under licence and their weights cannot leave the building, the safety research the paper calls for will be done by the same organisations it wants regulated, or by auditors with the access it demands. Either way the independent researcher is on the outside.

How far the labs got

We will be honest that the answer this week is not far. At least one lab has published a responsible scaling policy in the last few months, which is a form of self reported governance. None has committed to a third of its budget. None has offered white box access to outside auditors from the start of a training run. Registration, licensing and liability remain proposals.

The value of the paper, then, is as a marker. It states in one place, with signatures attached, what a specific group of senior researchers thinks would be adequate. That makes it possible to come back in a year and measure the distance between what was asked for and what was done. We intend to, and we would encourage anyone who cites the paper as a consensus to note who signed it, who did not, and which of its demands has moved.

Sources

  1. Bengio et al., Managing extreme AI risks amid rapid progress (arXiv 2310.17688)