DeepSeek open-infra week: the plumbing behind cheap MoE serving
Five days of repositories and one systems write-up showed that most of the DeepSeek cost story lives in kernels, all-to-all communication and disaggregated serving. Notes for people who run models rather than train them.
What was released, in order
Starting Monday, February 24, DeepSeek published one repository a day for five days, then a sixth post on the weekend describing how they actually serve V3 and R1. Day one was FlashMLA, a decoding kernel for their multi-head latent attention on Hopper GPUs. Day two was DeepEP, a communication library for expert parallelism. Day three was DeepGEMM, an FP8 matrix multiply library whose core logic is about 300 lines. Day four was DualPipe and EPLB, a pipeline schedule and an expert load balancer for training. Day five was 3FS, a distributed file system, with a data framework called Smallpond on top.
The training pieces got the attention, but the inference pieces are what changed our view of the January pricing. The R1 API price looked implausible to a lot of people who serve models for a living. After reading the day two and day six material, the price looks like the output of a specific set of engineering decisions rather than a subsidy. Most of those decisions are now public.
FlashMLA and the KV cache
MLA compresses the key and value projections into a shared latent vector, which is why V3 can hold long contexts in far less memory than a standard multi-head model of similar size. The catch is that the compressed cache needs a custom attention kernel to be fast. FlashMLA is that kernel. The launch README reports 3000 GB/s in memory-bound settings and 580 BF16 TFLOPS in compute-bound settings on an H800, with a paged KV cache using a block size of 64.
For someone running models, the point is that the architecture and the kernel are one decision, not two. If you take an MLA model and serve it with a generic attention implementation, you keep the memory saving but lose most of the speed. The open kernel is what lets other serving stacks catch up.
DeepEP and the all-to-all problem
A mixture-of-experts layer has to route every token to the experts it selects, which live on different GPUs, then gather the results back. That is two all-to-all exchanges per layer, and V3 has a lot of layers. DeepEP provides two families of kernels for this. The normal kernels are for training and prefill, where throughput matters. On a single node over NVLink the README reports 153 GB/s for dispatch and 158 GB/s for combine with eight-way expert parallelism. Across nodes over RDMA it lands around 43 to 47 GB/s, on ConnectX-7 InfiniBand cards rated at 400 Gb/s, which is roughly 50 GB/s of usable bandwidth. So the internode path is close to saturating the wire.
The low-latency kernels are for decoding, where each step moves a tiny amount of data and the round trip is what hurts. The README gives dispatch latency from 163 microseconds at EP 8 to 194 microseconds at EP 256, and combine from 318 to 369 microseconds over the same range. The library also offers hook-based overlap, so the RDMA transfer proceeds while the GPU computes without any streaming multiprocessors being reserved for communication. Dispatch runs in FP8 and combine in BF16, which halves the bytes on the wire in the direction that carries the most.
One detail we appreciated is the honesty about tricks. The README says the kernels use PTX instructions that are documented as undefined behaviour, chosen for speed and tested on Hopper specifically. That is the kind of thing that gets buried in a company's internal wiki. Putting it in the public README is a warning to anyone porting the code to other hardware.
How they actually serve it
The day six post is the piece we would hand to a colleague who runs inference. Prefill and decode are run on separate clusters. Prefill uses expert parallelism of 32 across four nodes, with each GPU holding nine routed experts plus the shared expert. Decode uses expert parallelism of 144 across 18 nodes, with two routed experts per GPU. The reason for the split is that the two phases want different batching and different communication patterns, and a single deployment forces a compromise on both.
Communication is hidden by splitting each batch into two microbatches and alternating them, so one microbatch computes while the other's all-to-all is in flight. Decoding uses a five-stage pipeline over a subdivided attention layer to get the same effect. Three separate load balancers handle prefill compute, decode KV cache occupancy, and expert imbalance across GPUs. The post reports 73.7 thousand input tokens per second per H800 node in prefill and 14.8 thousand output tokens per second per node in decode.
The economics are given for one 24 hour window. Peak occupancy was 278 nodes, average 226.75, at an assumed rental price of two dollars per GPU hour, for a daily cost of about 87,072 dollars. The system processed 608 billion input tokens, of which 342 billion hit the prefix cache, and generated 168 billion output tokens. At R1 list prices that would be 562,027 dollars of revenue, a theoretical margin of 545 percent. The post is careful to say actual revenue is much lower because V3 is cheaper, the web and app are free, and there are night discounts. We read the 545 figure as an upper bound on the hardware side of the story, not a claim about the business.
The engine that will not be released as an engine
After the week, DeepSeek posted a note about their inference engine itself. They want to open it, but it is a fork of vLLM from more than a year ago with heavy model-specific changes, it is tied to their cluster management, and they say plainly that a small team building models does not have the bandwidth to maintain a large open project. The plan is to extract standalone pieces as libraries and to contribute design details and implementation directly to existing engines such as vLLM.
We think that is the right call and also a useful signal. The cost advantage is not in any single artifact. It is in the coupling between architecture, kernels, communication library and cluster layout, and the week gave the rest of us the parts without the assembly instructions. What we would want someone to try next is a straight reproduction of the day six numbers on rented H800 or H100 nodes using only the published components, so we know how much of the margin is transferable and how much lives in the parts that stayed inside.
Sources
From the foundation