The numbers that matter are on the efficiency side

DeepSeek-V2 has 236B total parameters, activates 21B per token, was pretrained on 8.1T tokens and supports a 128K context. Those are respectable numbers. The numbers we keep coming back to are on the efficiency side. Relative to DeepSeek's own dense 67B model, training costs 42.5 percent less, the KV cache is 93.3 percent smaller, and maximum generation throughput is 5.76 times higher. The paper's thesis is that the KV cache is what caps how many users you can serve from a GPU, and the architecture is built around that.

The cost figures are stated in GPU hours on the lab's H800 cluster. Each trillion training tokens takes 300.6K GPU hours for DeepSeek 67B and 172.8K for DeepSeek-V2. On a single node of eight H800s the new model serves more than 50K output tokens per second and takes in prompts at over 100K tokens per second. Those are the kind of figures you publish when the point of the paper is that you can afford to keep doing this.

Why the KV cache is the bottleneck

During generation a transformer caches the key and value vectors of every previous token at every layer so it does not recompute them. Standard multi-head attention stores 2 times the number of heads times the head dimension times the number of layers elements per token. At long contexts and large batches this cache, not the model weights, fills the GPU memory, and it is what limits batch size and therefore throughput. The existing answers were grouped-query attention, which shares keys and values across groups of heads, and multi-query attention, which shares a single key and value across all heads. The paper's appendix ablation on 7B dense models says what practitioners already suspected, that GQA and MQA lose accuracy against MHA on hard benchmarks.

How multi-head latent attention works

MLA takes a different route. Instead of sharing keys and values across heads, it compresses them. For each token a single down-projection produces a latent vector of dimension d_c, and the per-head keys and values are recovered from it by up-projections. Only the latent is cached, so the cache per token is d_c times the number of layers rather than 2 times heads times head dimension. Because the up-projection for keys can be absorbed into the query projection at inference time, the model never has to materialise the full keys at all. The paper also applies a low-rank compression to queries, which does not shrink the cache but reduces activation memory during training.

The catch is rotary position embeddings. RoPE multiplies keys and queries by position-dependent matrices, and if the key up-projection sits between the latent and RoPE, it can no longer be absorbed. DeepSeek's answer is a decoupled RoPE: a small extra set of query heads and one shared key carry the positional information, and those are cached alongside the latent. The final cache per token is d_c plus the decoupled key dimension, times layers. With d_c set to four times the head dimension and the decoupled key at half a head dimension, the paper puts the total at the equivalent of GQA with 2.25 groups, while its ablations show MLA outperforming MHA.

Table 9 in the appendix gives the direct comparison. A large MoE at about 250B parameters with MHA needs 860.2K elements of cache per token and scores 57.5 on MMLU. The same model with MLA needs 34.6K elements per token and scores 59.0. On BBH the MLA variant leads 50.7 to 46.6. That is a 25-fold reduction in cache with no accuracy penalty in their setup, which is not how these trade-offs usually go.

DeepSeekMoE: fine-grained experts and shared experts

The feed-forward side is the DeepSeekMoE design the lab published in January, now at scale. Each MoE layer has 160 routed experts and 2 shared experts that every token passes through, with 6 routed experts activated per token. The argument for fine-grained experts is that splitting each expert into smaller pieces and activating more of them lets the router combine knowledge more precisely. The shared experts exist to hold common knowledge so that the routed ones do not all have to learn it.

The engineering detail that most reveals the constraint they were under is device-limited routing. The router is restricted so that each token's target experts span at most a fixed number of devices, which caps all-to-all communication cost under 8-way expert parallelism. There are auxiliary losses for balancing experts, devices and communication, and shared-expert computation is overlapped with the expert-parallel all-to-all. Batch size is ramped from 2,304 to 9,216 sequences over the first 225B tokens. This is a paper written by people who spent a long time watching interconnect utilisation.

Why this paper and not a later one

Everything that will make people take notice of this lab later is already here. A custom attention mechanism motivated by serving cost rather than benchmark score, a mixture-of-experts layout tuned to the communication topology of restricted hardware, training costs reported in GPU hours, and an open release with the numbers attached. The H800 is the export-controlled variant of the H100, and the paper reads as a systems-level answer to having less bandwidth to spend.

What we would want to see checked independently is the MLA ablation at the largest scale, since the appendix comparison is on their own training runs. We would also like someone to measure the real-world throughput gain at long contexts on hardware other than H800s, because a 93 percent cache reduction should translate into very different batch sizes at 128K. If those hold up, the KV cache conversation shifts from paged memory management, which PagedAttention solved for the cache you have, to making the cache small enough that it stops being the constraint.

Sources

  1. arXiv: DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model