What the letter asks for

On June 4 a group of current and former employees of frontier AI labs published a short open letter under the title A Right to Warn. Seven signed by name: Jacob Hilton, Daniel Kokotajlo, Ramana Kumar, Neel Nanda, William Saunders, Carroll Wainwright and Daniel Ziegler. Six more signed anonymously, four of them current employees. The signatories come from OpenAI, Google DeepMind and Anthropic, and the letter is endorsed by Yoshua Bengio, Geoffrey Hinton and Stuart Russell.

The letter makes four requests of what it calls advanced AI companies. That they not enter into or enforce agreements that prohibit disparagement or criticism over risk concerns, and not retaliate against people who raise them. That they set up a verifiably anonymous process by which current and former employees can raise risk concerns to the board, to regulators, and to an appropriate independent organisation. That they support a culture of open criticism in which employees can raise risk concerns publicly, subject to protecting trade secrets. And that they not retaliate against employees who share risk-related confidential information publicly after other processes have failed, with the caveat that people should go through the anonymous channel first where one exists.

The argument underneath is compact. The companies hold substantial non-public information about what their systems can do and what safeguards exist. They have financial incentives to avoid oversight and there is little government oversight to speak of. Ordinary whistleblower protections cover illegal conduct, and most of what the signatories worry about is not yet regulated, so it is not illegal. That leaves employees as the main line of accountability, and employees are bound by broad confidentiality agreements that, the letter says, can prevent them from voicing concerns except to the very companies that may be failing to address them.

Why this is a research culture problem

We read this letter as a document about how research works inside closed labs, more than as a document about catastrophic risk. The scientific method depends on a specific kind of speech: someone saying, in public, that a result does not hold or that a method has a failure mode. Academic labs make that speech cheap. You can criticise your advisor's paper at a conference and nothing happens to your visa. Industrial labs make it expensive, and the letter is a description of how expensive it has become at the frontier.

The tell is the first request. A non-disparagement clause is not a trade secret protection. Trade secrets are already protected by the confidentiality agreements everyone signs. A non-disparagement clause protects reputation, and a lab that needs its former researchers contractually bound not to say the safety work is inadequate is a lab whose research claims cannot be checked by the people best placed to check them. That is a validity problem for every safety result the lab publishes, whatever the truth of any individual claim.

The second request, an anonymous channel to the board, is more interesting than it looks. Boards of these companies have recently been shown to have limited visibility into what the executives know. An anonymous channel from staff to directors is a mechanism for correcting that, and it is telling that the people asking for it are the staff rather than the directors.

What outsiders can actually do

It is easy to read a letter like this and conclude that everything depends on the labs' goodwill. We think that is wrong in three concrete ways, and we want to lay them out because each is something a foundation like ours or a university group can act on now.

First, replicate what can be replicated. Every safety claim in a system card that rests on an evaluation with a public description can be re-run on the model's public API by someone with no confidentiality agreement. Where the numbers match, good. Where they do not, that discrepancy is information the labs' own staff cannot publish and we can. External replication is the one form of criticism no non-disparagement clause reaches.

Second, build the independent organisation the letter's second request refers to. The signatories ask for a channel to an appropriate independent body with relevant expertise, and no such body exists in a form a lab could route reports to today. Standing one up requires the technical capacity to assess a report and the legal structure to protect the reporter. Both are within reach of the academic and non-profit community, and neither requires a lab's permission.

Third, make the contracts legible. Whether a given lab's separation agreement includes a non-disparagement clause, and whether it has been waived, is a fact that can be asked about directly and reported. A clause that a lab is willing to defend in public is a different thing from one that only survives because nobody has asked.

What we would not claim

The letter does not allege that any specific safety failure has been covered up, and we are not going to infer one. Thirteen people out of thousands signed. Most lab researchers we know think their employers are behaving reasonably, and they may be right. The letter's strength is that it does not depend on any particular scandal. It asks for the conditions under which a scandal, if one occurred, could come to light. That is a modest request, and the fact that it needed a public letter and endorsements from three of the field's most senior figures to make it is the most revealing thing about the current state of the field.

What we would like to see next is a lab publishing its whistleblower policy in full, with the anonymous channel's design, the independent recipient named, and the retaliation protections spelled out. The first one to do so sets the standard the others will be measured against.

Sources

  1. A Right to Warn about Advanced Artificial Intelligence (open letter)