Are the labs losing money on inference? Doing the arithmetic
A widely shared napkin estimate this week argues that serving a frontier-sized model at public API prices leaves gross margins above 80 percent. We rework the numbers and walk through the assumptions that swing the answer, above all utilisation, context length and caching.
The estimate
Martin Alderson posted a back-of-envelope calculation on August 27 that has been passed around for three days now. He is explicit that it is napkin math, that it covers raw compute only, and that he has not run a frontier model at scale himself. The model he costs is DeepSeek R1, chosen because its architecture is public: 671 billion parameters with 37 billion active per token through a mixture of experts. The hardware is a cluster of 72 H100s rented at 2 dollars an hour each, which he notes is above current on-demand retail, for 144 dollars an hour in total.
The cluster is split into nine instances of eight GPUs each with tensor parallelism. Each instance serves a batch of 32 concurrent requests, and the sequences are assumed to average about 1,000 tokens. Everything downstream follows from those choices, so it is worth being clear that they are choices.
Prefill is nearly free, decode is not
The core of the argument is that a decode step is bound by memory bandwidth. Each step has to read the active weights, about 74 gigabytes in FP16 for 37 billion parameters, and an H100 moves 3.35 terabytes a second from HBM. Dividing gives roughly 45 forward passes a second per instance. In prefill, one forward pass processes every input token in the batch at once, so 45 passes a second over 32 sequences of 1,000 tokens comes to about 13 million input tokens a second across the nine instances, or 46.8 billion an hour. He takes 30 to 50 percent off for MoE routing inefficiency and the number is still enormous.
Decode generates one token per sequence per pass, so the same 45 passes a second yield 1,440 output tokens a second per instance and about 12,960 across the cluster, or 46.7 million an hour. Put 144 dollars an hour over those two throughputs and the cost of input comes out around a tenth of a cent per million tokens, while output costs around 3 dollars per million. That is the thousand-fold asymmetry the post is built on, and it is the reason the conclusion is what it is: at public list prices of 3 to 15 dollars per million output tokens, he estimates gross margins of 80 to 95 percent, with subscription plans marked up by anywhere from five to twenty times their compute cost.
Reworking it: the assumptions that matter
We redid the arithmetic with the same public figures and it holds up as arithmetic. What decides the answer is four assumptions, and each can move the result by a large factor. The first is utilisation. The estimate assumes every instance runs a full batch of 32 all day. Real traffic is bursty, and a cluster provisioned for the peak sits partly idle the rest of the time. If average utilisation is 40 percent, the per-token cost rises by two and a half times before anything else changes.
The second is batch size itself. Thirty-two concurrent sequences per instance is achievable when contexts are short, because the KV cache for each one is small. At long contexts the cache grows, memory fills, and the batch has to shrink, which cuts output throughput in proportion. The third is context length on the compute side. Alderson flags this himself: at 128,000 tokens and beyond, attention becomes the dominant cost and the workload shifts from memory-bound to compute-bound, which he puts at two to ten times more expensive. The sequences in the estimate average 1,000 tokens. The agent traffic that now dominates spend averages far more.
The fourth is everything left out. The post costs GPUs only, with no networking, storage, staff, redundancy or failed experiments, and it uses retail H100 pricing that large buyers undercut. Those pull in opposite directions, and we would not try to net them out from the outside.
Caching changes the sign of the input assumption
The post does not discuss prompt caching, and we think it is the biggest missing term for the workloads that matter. A coding agent sends the same large context, the repository state and the conversation so far, on every turn with a small addition at the end. Without caching, that context is re-prefilled every turn, and while prefill is cheap per token, it is not cheap when the context is 100,000 tokens and the turn produces 200. With caching, the KV state for the shared prefix is stored and only the new tokens are computed, so the effective input cost collapses again.
Which way this moves the margin depends on what the provider charges for cached tokens against what storing the cache costs in HBM or in a tier below it. The thousand-fold asymmetry in the post is a statement about compute. Once caching is in the picture the binding constraint on input can become memory, and we have not seen a public estimate that prices that side properly.
What we take from it
The estimate's conclusion is that raw inference is well within the price, and we think that survives the rework. The refined version is that margins are large on short-context chat and much thinner, possibly negative in bad cases, on long-context agent traffic at high concurrency, and that the difference is mostly about whether the KV cache fits. It follows that pricing should diverge by workload, and that the labs with the cheapest long-context serving will be the ones with the most room to cut agent prices.
What we would like someone to publish is the same calculation with a realistic distribution of context lengths taken from actual agent logs rather than a single 1,000-token average. That one change probably does more to the answer than every other assumption combined.
Sources
From the foundation