The sentence that stopped us

Mark Zuckerberg sat down with Dwarkesh Patel on April 18, the day Llama 3 shipped, and somewhere in the middle of a conversation about model sizes he said that no one has built a one gigawatt datacenter yet. He described the current range as facilities of 50 to 150 megawatts, with some at 300 to 500, and said a gigawatt site would be the size of a meaningful nuclear power plant devoted to training models. He was not describing a hypothetical. He was describing the thing he expects to need.

We have spent years thinking about compute in units of GPUs and GPU hours. Hearing the head of one of the four or five organisations that can actually run frontier training talk in units of power plants was a reframing, and we want to write down why before it becomes ordinary.

Chips were the bottleneck last year

The interview is full of the numbers people expected. Llama 3 launched with an 8B and a 70B model, the 70B trained on around 15 trillion tokens, which Zuckerberg said was well past what a compute-optimal recipe would call for and that the model was still improving when they stopped. A 405B dense model was still training and already around 85 on MMLU. Meta has two clusters of roughly 22,000 to 24,000 GPUs for large training runs and expects about 350,000 H100s across its fleet by the end of the year.

What surprised Patel, and us, was where Zuckerberg put the constraint. He said the problem with building the next size of datacenter is not primarily capital. It is that energy is a heavily regulated government function, with permitting, transmission line approvals and land decisions that carry years of lead time. Nvidia's allocation was the story of 2023. On this account the story of the next few years is utilities and regulators, and a company with tens of billions to spend cannot buy its way past a multi-year interconnect queue.

He also made a point about inference that we think gets lost when people count training FLOPs. Meta's compute goes overwhelmingly to serving billions of users, and a large share of the GPU fleet runs ranking models for Reels and the feeds rather than language models at all. The 350,000 H100 figure is a fleet number, not a training number, and the gap between those two is the shape of a consumer company rather than a lab.

Open weights as a distribution of power

The other half of the interview was the argument for releasing weights. Zuckerberg's framing was about concentration. He said the scenario that worries him more than widespread access is an untrustworthy actor holding the single strongest model, and he compared open weights to the way open source software keeps any one entity from dominating security. He also gave the economic reason, which is that community improvements to an open model save Meta money at Meta's scale, with the Open Compute Project as the precedent he keeps returning to.

There were two qualifications that we think matter more than the headline. The Llama licence restricts the very largest companies from reselling the models without a separate deal, so this is open with a carve-out for direct competitors. And when Patel pushed on whether Meta would keep releasing at every capability level, Zuckerberg said that if at some point there is a qualitative change in what the thing can do and releasing it seems irresponsible, they will not. He declined to name a threshold and said it would be judged in real time.

Put those together and the energy point becomes part of the same argument. If the next generation of models needs a gigawatt of dedicated power and years of permitting, the number of organisations that can train one shrinks to those that can build utility-scale infrastructure. Open weights are a way of saying that the output of that infrastructure, if not the infrastructure itself, is not a moat. Whether that holds when the release decision is made in real time by one company is exactly the question Patel was asking.

What this means for a lab like ours

Nothing in this interview changes what a small foundation can do in a training run. It changes the picture of what we are downstream of. When the input to frontier work is measured in power plants, the research that matters at our scale is the research that makes more of a given model, and the research that checks whether the models being handed to us behave the way their release notes say. Both of those depend on weights being available, and the interview was a clear statement that availability is a policy that can be reversed.

The empirical question we would like to see tracked is simple. Every quarter, list the training sites that exist above 300 megawatts and who owns them. If that list stays at a handful of names for the next three years, then the energy constraint has done what export controls on chips were meant to do, and the open weights conversation is really a conversation about the goodwill of five companies. We would rather know that from a table than from an interview.

Sources

  1. Dwarkesh Patel: Mark Zuckerberg on Llama 3, open sourcing $10b models, and energy (April 18, 2024)