What was promised

In July 2023 OpenAI announced a Superalignment team, co-led by Ilya Sutskever and Jan Leike, with the goal of solving the alignment of superintelligent systems within four years. The commitment that got the most attention was the resourcing line. OpenAI said it would dedicate 20 percent of the compute it had secured to date, over the next four years, to the effort. For a company whose main constraint is GPUs, that was a large number, and it was read at the time as a signal that safety research would have a claim on the same resources as product.

Ten months on, the team is gone. Sutskever's departure was announced last week. Leike resigned on Friday, May 17, and said in public that over the past years safety culture and processes had taken a backseat to shiny products, and that he had gradually lost trust in the company's leadership. The roughly 25 remaining members of the team have been reassigned within the company. Leike has since joined Anthropic.

What was delivered

It would be unfair to say the team produced nothing. The weak-to-strong generalisation paper, posted in December 2023 with Burns, Izmailov and ten coauthors including Leike and Sutskever, is a genuine contribution. It asked whether a weak supervisor can elicit the full capability of a much stronger student, using GPT-2-level models to supervise GPT-4, and found that with an auxiliary confidence loss the student recovered performance close to GPT-3.5 level on NLP tasks. That is a real experimental handle on a problem that had mostly been discussed in the abstract, and it is the kind of work the team was set up to do.

But one paper in ten months is not what 20 percent of a frontier lab's compute buys. Fortune's reporting, based on six sources, says the team never received anywhere close to that share. It got its compute through the ordinary quarterly budget process, at a far lower level, and its requests for additional flexible GPU capacity were routinely turned down by leadership. One source noted that the pledge itself was never pinned down: 20 percent of compute secured as of July 2023, spread over four years, could mean 20 percent every year or something much smaller front-loaded or back-loaded, and nobody outside the company could tell which reading was in force.

The audit problem

That last detail is the one we keep coming back to. The pledge was phrased in a way that sounded precise and was in fact unverifiable. There was no baseline figure for compute secured to date, no reporting cadence, no external party who could confirm the allocation, and no definition of what counted as superalignment work versus adjacent research. A commitment with those properties can be honoured or ignored without anyone outside the building being able to tell the difference, and in this case the people inside the building are telling us it was ignored.

Leike's own framing supports that. He wrote that for the past few months his team had been sailing against the wind, sometimes struggling for compute, and that it was getting harder and harder to get the research done. Read alongside the original announcement, that is a description of a resourcing commitment that quietly lapsed without any public revision. The company did not announce that the 20 percent figure had been reduced. It simply did not deliver it, and the world found out when the people who needed it left.

What a checkable commitment would look like

We do not think the lesson is that labs should stop making public safety commitments. The lesson is that a commitment only constrains behaviour if a third party can check it. For compute, that would mean publishing the baseline the percentage is measured against, reporting the allocation at fixed intervals in units that cannot be reinterpreted, and letting an outside body, an auditor, a regulator, or a named academic group, verify the reports against internal accounting. None of that is technically hard. It is uncomfortable, which is a different thing.

The same standard should apply to the research output. A team funded at a fifth of a frontier lab's compute should be producing results at a rate that reflects it, and the gap between the weak-to-strong paper and everything else on the ledger should have been visible to outsiders long before the resignations. Publication is the cheapest audit there is. When a well-funded team goes quiet, that silence is data.

What we would want next

Every frontier lab now has some version of a safety commitment on its website. The useful exercise this week is to go through each one and ask, for every number and every promise, who outside the company could tell if it was broken. Where the answer is nobody, the promise is a press release. We would like to see the next lab that makes a compute pledge attach a reporting mechanism to it on day one, and we would like the research community to treat pledges without one as unverified claims, which is what they are.

Sources

  1. Fortune, OpenAI promised 20 percent of its computing power to combat the most dangerous kind of AI, but never delivered
  2. Wikipedia, Jan Leike
  3. Wikipedia, Superalignment
  4. Burns et al., Weak-to-Strong Generalization (arXiv 2312.09390)