SB 1047 vetoed: the frontier-model bill that split the safety community
Newsom returned SB 1047 unsigned on September 29, arguing that a bill keyed to training cost and compute ignores where a model is deployed. A look at what the bill required, who lined up on each side, and why the veto message reads like a research question.
What the bill would have done
SB 1047 applied to models that cost more than 100 million dollars to train and used more than 10 to the 26 floating point operations, plus fine-tunes that cost more than 10 million dollars. Developers of those models would have had to write a safety and security protocol before training, file a statement that they had taken reasonable care to prevent critical harms, and submit to annual third-party audits from 2026. Critical harms were defined narrowly: chemical, biological, radiological or nuclear weapons, cyberattacks on critical infrastructure causing mass casualties or 500 million dollars in damage, autonomous crimes of the same scale, and comparable harms.
The bill also carried a shutdown capability requirement, whistleblower protections for employees, and CalCompute, a public compute cluster run through the University of California for researchers and startups. A nine-member Board of Frontier Models would have overseen the whole thing. The version that reached the governor was already softer than the one introduced on February 7. Amendments on August 15 removed the proposed Frontier Model Division, dropped a perjury penalty, changed the standard from reasonable assurance to reasonable care, and narrowed the shutdown requirement for open-weight developers.
The votes were lopsided. The Senate passed it 32 to 1 in May, the Assembly 48 to 16 on August 28, and the Senate again 30 to 9 on August 29 after the amendments. The division was among the people who build and study these systems.
Who was on which side
The supporter list was heavy with people who have spent years on catastrophic risk: Yoshua Bengio, Geoffrey Hinton, Stuart Russell, Dan Hendrycks, Jan Leike, Max Tegmark, Kevin Esvelt. More than 113 current and former employees of OpenAI, Google DeepMind, Anthropic, Meta and xAI signed in support, along with the OpenAI whistleblowers Daniel Kokotajlo and William Saunders. Elon Musk backed it. So did SAG-AFTRA and the Los Angeles Times editorial board. Anthropic, after its amendments were adopted, wrote that the benefits likely outweighed the costs while adding that it was not certain and that some aspects still seemed concerning or ambiguous.
The opposition was just as credentialed and drew from a different tradition. Fei-Fei Li warned the bill would harm the American ecosystem. Andrew Ng argued for targeting specific harms such as deepfake pornography and watermarking instead of model size. Yann LeCun said it would kill open source models. Ion Stoica and Jeremy Howard opposed it, as did researchers at the University of California and Caltech in open letters. OpenAI and Meta raised objections. Y Combinator and Andreessen Horowitz lobbied against it, and so did Nancy Pelosi, Zoe Lofgren, Ro Khanna and six other members of California's congressional delegation.
The open-weights argument was the sharpest edge. The AI Alliance and others worried that a developer like Meta would stop releasing Llama weights if it could be held liable for what a downstream user did. Lawrence Lessig argued the opposite, that the bill would make open models safer and more attractive to developers. Both sides were predicting the behaviour of a handful of companies under a law that had never existed, which is why the argument never resolved.
What the veto message actually says
The message is two and a half pages and worth reading in full, because it does not say what either camp expected. Newsom does not say the risks are imaginary. He writes that he agrees with the author that we cannot afford to wait for a major catastrophe before acting, and that safety protocols must be adopted. He disagrees on the trigger. The key question, in his words, is whether the threshold for regulation should be based on the cost and number of computations needed to develop a model, or whether we should evaluate the system's actual risks regardless of those factors.
His answer is that a compute threshold gets the risk model wrong in both directions. By focusing only on the most expensive and large-scale models, the bill could give the public a false sense of security, while smaller, specialized models may emerge as equally or even more dangerous. And in the other direction, the bill does not take into account whether an AI system is deployed in high-risk environments, involves critical decision-making or the use of sensitive data. Instead it applies stringent standards to even the most basic functions, so long as a large system deploys it.
He then asks for something the bill did not have: a solution informed by an empirical trajectory analysis of AI systems and capabilities. The accompanying announcement names Fei-Fei Li, Tino Cuéllar and Jennifer Tour Chayes to lead that work, directs Cal OES to assess generative AI risks to energy, water and communications infrastructure, and notes that he signed 17 narrower AI bills in the preceding 30 days covering deepfakes, watermarking and election misinformation.
The technical question hiding in the political one
Strip out the politics and the veto poses a question our field has not answered. Is dangerous capability a function of training compute, or of deployment context, or of something we cannot yet measure? SB 1047 bet on compute because compute is observable before a model exists and everything else is observable only afterward. Newsom's objection is that the bet is wrong at both ends: a large model doing autocomplete is regulated, a small fine-tune advising on pathogen synthesis is not.
We think both positions are half right and the reason is that we lack the evidence the message asks for. Nobody has a validated way to predict from a training run which capabilities will appear, and nobody has a validated way to say which deployment contexts turn a capability into a harm. A compute threshold is a crude proxy adopted because it is the only thing that can be measured in advance. A deployment-context rule needs a taxonomy of contexts and harms that does not exist in usable form.
What we would want to see next is the trajectory analysis treated as a research programme with published methods, rather than a panel report. Which capability evaluations, run on which models at which compute scales, actually predicted later real-world incidents? Until someone can answer that with data, every bill on this subject will be an argument between two guesses, and the governor will be right that neither guess deserves the force of law.
Sources
From the foundation