Blackwell and the year compute became a policy variable
Nvidia's GB200 NVL72 promises up to 30 times the inference throughput of H100 and every major cloud has already signed on. What a single vendor's roadmap means for who gets to do frontier research, and how a small non-profit should plan underneath it.
What was announced
Nvidia used its GTC keynote on March 18 to announce the Blackwell platform. The GPU itself has 208 billion transistors on a custom TSMC 4NP process, with a 10 terabyte per second link between its two dies and 1.8 terabytes per second of bidirectional NVLink bandwidth per GPU. The product most of the announcement is about is the GB200 NVL72, a liquid-cooled rack that connects 36 Grace Blackwell superchips, meaning 72 Blackwell GPUs and 36 Grace CPUs, into a single system with 30 terabytes of memory and a claimed 1.4 exaflops of AI performance. Fifth-generation NVLink can tie up to 576 GPUs together.
The headline comparison is against H100. For large language model inference Nvidia claims up to 30 times the performance, and up to 25 times lower cost and energy use. Those are vendor numbers for a favourable workload, and we would treat them as upper bounds until independent measurements appear. Even at a third of the claim, it is the largest single-generation jump in serving capacity the field has seen.
Everyone signed the same page
The part of the press release that we keep rereading is the list of names. AWS, Google Cloud, Microsoft Azure and Oracle are all committed to offering Blackwell instances. So are CoreWeave, Lambda, Crusoe, Nebius, Applied Digital and IBM Cloud. On the customer side the release quotes OpenAI, Meta, Tesla and xAI, with Elon Musk saying there is currently nothing better than Nvidia hardware for AI. Every server manufacturer that matters, from Dell and Supermicro to Foxconn and Wiwynn, is listed as building systems around it.
That is not a market with several roadmaps competing. It is a market with one roadmap and a queue. When a single vendor's release cadence determines when every lab in the world gets its next capacity step, that cadence becomes something closer to a policy variable than a product decision. Export controls already treat Nvidia part numbers as the unit of regulation. Allocation between customers, pricing and delivery order are now among the most consequential decisions about who can run what, and they are made in one company.
What it does to frontier research
The gap between labs that get early Blackwell allocation and those that do not is going to be measured in what experiments are affordable. A 30 times inference improvement, if it holds even partly, changes which methods are reasonable. Anything that spends most of its compute at sampling time, from search-heavy reasoning to large-scale synthetic data generation to running many agent trajectories for evaluation, becomes much cheaper for whoever has the hardware first. Methods that were curiosities on H100 become defaults on GB200.
The consequence for the rest of us is that the frontier moves in step with delivery schedules we do not control. A university group or a small foundation is not going to be in the first wave of NVL72 racks. We will read papers whose experiments assume a cost per token that we cannot match for a year or more, and we will have to decide which of those results we can reproduce at reduced scale and which we simply have to take on trust. That is uncomfortable for a foundation whose whole purpose is that results should be checkable.
A compute strategy for a small non-profit
Here is how we think about our own position underneath this. First, do not compete on scale. There is no version of our budget that buys a rack of GB200s in the first year, and pretending otherwise would waste money on last-generation hardware at peak prices. Second, rent rather than own for anything bursty. The cloud providers in the announcement will all be offering Blackwell, and the competition between the neoclouds on the list should keep hourly rates moving in our favour over time. Owning hardware makes sense for steady workloads, and most of our workloads are not steady.
Third, pick problems where the interesting quantity is a curve rather than a point. A scaling study across small models tells you something a single frontier run does not, and it can be done on hardware two generations old. Fourth, invest in evaluation and reproduction, because those are the things the field will need more of as the frontier gets more expensive, and they scale down more gracefully than training does. A careful reproduction of a published result at a tenth the scale, with all the code released, is worth more to the community than a poorly resourced attempt to match the original.
What we will be watching
The number we want is the real delivered cost per token for inference on GB200 versus H100, measured by a customer rather than the vendor, across a few model sizes and batch settings. The second thing we want is allocation transparency. Which labs got racks in which quarter is a fact that shapes research output, and right now it is known only through leaks. If compute is going to be a policy variable, the least the field can ask is that the variable be observable.
Sources
From the foundation