A threshold with an exponent

Until Monday, every government statement about frontier AI described it in adjectives. Executive Order 14110, signed on October 30, describes it in floating point operations. A dual use foundation model that must be reported to the Commerce Department is one trained with more than 10 to the 26 integer or floating point operations, or more than 10 to the 23 if it was trained primarily on biological sequence data. A computing cluster that must be reported is one in a single datacenter, connected by networking of over 100 gigabits per second, with a theoretical peak of 10 to the 20 operations per second.

Those figures are interim. The order directs Commerce to update them, and the ninety day clock on defining the reporting requirements started on Monday. But they are the first time a regulator has written down a compute number as the line between a model that is nobody's business and a model the government wants to hear about, and the exponent will now be argued over the way every other regulatory threshold is.

What the reporting actually requires

The legal lever is the Defense Production Act. Section 4.2 uses it to require companies developing models over the threshold to report, on an ongoing basis, their training activities, who owns and how they protect the model weights, and the results of red team testing. Until NIST publishes standards, the interim red team reports must cover biological weapons, discovery of software vulnerabilities, influence operations, and the possibility of self replication. NIST has 270 days to produce the guidelines and red teaming standards.

There is a separate track for cloud providers. Commerce is to propose rules within 90 days requiring infrastructure as a service providers to report when a foreign person uses their compute to train a model over the threshold, and within 180 days to propose identity verification for foreign customers. The order does not prohibit anything about model training. It builds a reporting pipeline and leaves what to do with the reports for later.

What twenty eight governments agreed to at Bletchley

The declaration signed on November 1 at Bletchley Park is a different kind of document. It has no thresholds and no deadlines. It has signatures from 28 countries plus the European Union, and the list includes both the United States and China, alongside Brazil, India, Indonesia, Kenya, Nigeria, Saudi Arabia, the UAE, Japan, South Korea and the large European economies. Getting Washington and Beijing on the same page about anything involving frontier models is the achievement, and the price of it is text that commits nobody to much.

What the text does do is name the concern precisely. It singles out highly capable general purpose models, including foundation models, and says they could cause serious, even catastrophic harm through misuse or through problems of control, with cybersecurity and biotechnology called out as the domains of particular worry. The agenda has two items, building a shared scientific understanding of frontier AI risks across countries, and building risk based national policies with transparency, evaluation metrics and safety testing tools. The signatories agreed to meet again in 2024.

Why the pairing matters

Read together the two documents divide the labour. The executive order is a domestic mechanism that produces information, red team results and weight security reports from the handful of companies that cross the compute line. The declaration is an international frame that says the information matters and that other governments want their own version. One is enforceable and narrow. The other is broad and voluntary. Neither would mean much without the other.

The thing we notice as a researcher is what both documents are asking for and do not yet have. Red team standards do not exist, which is why NIST got 270 days. Evaluation metrics for catastrophic misuse do not exist, which is why the declaration calls for building them. The week produced a demand for a science of dangerous capability evaluation, and that science is currently a few dozen people and a handful of preliminary papers.

A note added two years on

Looking back from late 2025, the two documents had different fates. Executive Order 14110 was rescinded on January 20, 2025, within hours of the presidential inauguration, as part of a package of rescissions of the prior administration's actions. The compute reporting regime it created was therefore short lived as a legal matter, though the 10 to the 26 figure had already spread into other jurisdictions' drafting and into the vocabulary of the field.

The summit series survived. The Bletchley meeting was followed by a summit in Seoul in 2024, one in Paris in 2025, and further meetings are planned for New Delhi in 2026 and Geneva in 2027. Whether the successor meetings kept the safety focus of the first one is a longer argument. What lasted from that week in November 2023 was the idea, now ordinary, that a frontier model is a thing you can define by its training compute and that governments should expect to be told when one is built.

Sources

  1. Executive Order 14110 on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence (October 30, 2023)
  2. The Bletchley Declaration by Countries Attending the AI Safety Summit, 1-2 November 2023
  3. Wikipedia, Executive Order 14110
  4. Wikipedia, AI Safety Summit