The arithmetic

David Cahn at Sequoia published a follow-up last week to a piece he wrote in September, and the number in the title tripled. The calculation is the same both times. Take Nvidia's data centre revenue run rate. Double it, on the assumption that a GPU is roughly half the cost of the data centre it lives in once you add energy, buildings and networking. Double it again, on the assumption that whoever buys the compute needs a 50 percent gross margin to make the purchase sensible. The result is the revenue the AI industry needs to generate to pay for the hardware it has bought.

In September that came to 200 billion dollars against a generous 75 billion of plausible revenue, for a gap of 125 billion. Nine months later the run rate has grown enough that the implied revenue requirement is around 600 billion. On the other side of the ledger the one real number is OpenAI, which Cahn puts at 3.4 billion dollars in revenue, up from 1.6 billion late last year. Everything else is an estimate of what large technology companies might be earning from AI features.

What changed since September

The piece lists the shifts since the first version. The GPU shortage that defined 2023 has eased, and hyperscalers are building stockpiles rather than fighting for allocation. Nvidia has announced the B100, which Cahn summarises as about 2.5 times the performance for about 25 percent more cost, and which will push buyers to wait for the next chip and will push down the resale value of the current one. He frames the depreciation risk as underappreciated by the market. A data centre full of H100s is not a long-lived asset if the chip that replaces it is two and a half times better a year later.

His conclusion is that GPU compute is heading toward a commodity, with the pricing power sitting with the chip vendor rather than with anyone who owns the chips, and that history is not kind to people who build speculative infrastructure ahead of demand. The recommendation to founders is the same as in September. Build things that create value for an end customer rather than betting on the infrastructure layer.

Reading it from a small lab

We do not own GPUs. We rent them by the hour from whoever has capacity, and we run experiments that would have been large a few years ago and are now small. That puts us in a specific position relative to Cahn's argument, which is that every mechanism he describes as a risk to the people who bought the hardware is a benefit to the people who rent it.

If the shortage has eased, our jobs schedule sooner. If stockpiles are growing, someone will rent them out below cost rather than leave them idle. If the B100 arrives at 2.5 times the performance, the H100 hourly rate falls, and most of our work does not need the newest chip. If a correction comes and the marginal data centre is sold off, the buyer will have paid less and will charge less. The 600 billion dollar gap is, from the cheap seats, a description of subsidised compute.

This is not schadenfreude. The same overbuild that makes our compute cheap is what trained the models we study, and if the capital dries up the frontier stops moving and there is less to study. We would rather the industry find the revenue. But the honest reading of the arithmetic is that a research group that does not need to own hardware is on the right side of a correction, and that it is worth planning as though hourly prices keep falling.

Where we think the argument is weakest

The 50 percent margin assumption does the most work and gets the least attention. It is the right assumption for a business that resells compute, and the wrong one for a business that uses compute internally, where the return shows up as products that would not otherwise exist rather than as a margin on GPU hours. Most of the capital in question belongs to companies in the second category. The gap does not disappear on that reading. It does mean the 600 billion figure is a requirement on an accounting basis that does not apply to most of the buyers.

The other weakness is the treatment of the revenue side as static. OpenAI's number doubled in six months. If it keeps doing that the gap closes on its own within a few years, and if it does not the gap is real. The interesting question is which, and the piece does not attempt to forecast it. It is a description of a mismatch at a point in time, and it is a good one, but the mismatch is between a stock of spending and a flow of revenue, and flows can grow.

What we will watch

The number we care about is the hourly rental price for an H100 from the second tier providers, because that is where surplus capacity shows up first. If Cahn is right it should fall through the autumn as B100 deliveries begin, and a lab like ours should be able to run the same experiments for less each quarter. If it holds, the demand is real and the gap is a timing problem rather than a bubble.

The second thing is whether the revenue estimates Cahn treats as speculative become real numbers. Only one company has published a figure. When the others do, the calculation can be redone with fewer assumptions, and we expect the answer to be more boring than either the bulls or the bears would like.

Sources

  1. David Cahn, AI's $600B Question (Sequoia Capital, June 20, 2024)
  2. David Cahn, AI's $200B Question (Sequoia Capital, September 20, 2023)