DeepSeek V4: 27 percent of the FLOPs per token
DeepSeek shipped V4-Pro and V4-Flash under MIT with a million tokens of context and a claim that Pro needs 27 percent of V3.2's per-token compute at long context. A trace through the efficiency stack and a question about what the phrase three to six months behind is measuring.
What shipped
DeepSeek released two models on April 24. V4-Pro has 1.6 trillion total parameters with 49 billion active per token. V4-Flash has 284 billion total with 13 billion active. Both are mixture-of-experts models with a one million token context window, both are pretrained on more than 32 trillion tokens, and both are released under the MIT licence with weights on Hugging Face, 865 gigabytes for Pro and 160 for Flash. By parameter count Pro is the largest open-weights model anyone has released, ahead of Kimi K2.6 at 1.1 trillion and GLM-5.1 at 754 billion.
API pricing is 14 cents per million input tokens and 28 cents output for Flash, and 1.74 dollars input and 3.48 output for Pro. The models ship with three modes, a non-thinking mode, a high reasoning mode, and a maximum mode that uses a special system prompt.
The efficiency claim
The number in the title comes from the paper's own statement that at one million tokens of context V4-Pro attains only 27 percent of the single-token inference FLOPs and 10 percent of the KV cache size of DeepSeek-V3.2. That is a claim about long-context cost specifically. At short context the ratio will be different, and it is worth keeping that qualifier attached when the figure gets repeated.
Three architectural pieces do the work. The first is the attention design, which the model card describes as a hybrid of Compressed Sparse Attention and Heavily Compressed Attention. Both are ways to avoid having every token attend to every other token at full resolution, and the ten-to-one reduction in KV cache is mostly their doing, since the cache is where long context spends its memory. The second is the mixture-of-experts routing that has been DeepSeek's signature for several generations, and which is why a 1.6 trillion parameter model can run at 49 billion active. The third is a change to the residual stream the card calls manifold-constrained hyper-connections, described as strengthening conventional residual connections to stabilise signal propagation across layers without giving up expressivity.
On top of that, training used the Muon optimizer, and post-training ran in two stages: cultivate domain-specific experts, then consolidate them into one model through on-policy distillation. We mention the training recipe because efficiency at inference is only part of the story. A model that is cheap to serve and expensive to train is a different bet from one that is cheap at both ends, and the report gives enough to see that DeepSeek is optimising both.
Where each piece comes from
None of these components appeared fully formed in V4. Sparse attention was the headline of V3.2, and the V4 version extends it. The expert routing has been refined over several generations. What is new is the combination at this scale and the willingness to state the efficiency numbers relative to the company's own previous model, which is a more checkable comparison than the usual one against a competitor's undisclosed architecture.
The honest caveat is that the 27 percent figure is DeepSeek's measurement of DeepSeek's models under DeepSeek's serving stack. It is plausible, and it is consistent with the pricing, but nobody outside has reproduced it yet. With MIT weights and the architecture described, someone will, and we would like to see the number at 8k and 128k context alongside the million-token one.
What three to six months behind measures
The line everyone quoted is from the paper itself. On the self-reported benchmarks, V4-Pro in its maximum mode beats GPT-5.2 and Gemini 3.0 Pro on standard reasoning tests and falls slightly short of GPT-5.4 and Gemini 3.1 Pro, from which the authors conclude they trail the frontier by roughly three to six months. We want to take that phrase apart, because it is doing more work than it looks.
Measured in benchmark score, the claim is that an open model today matches a closed model from a few months ago. That is a statement about a scalar, and scalars hide things. It does not say whether the gap is on the same tasks or different ones, whether it closes further under a fixed compute budget, or how the models compare on anything not in the table. Measured in cost, the picture inverts. A model you can download, run and fine-tune at 27 percent of last year's compute sits on a different axis from the one the phrase describes.
The phrase is also a claim about a rate. Behind by three to six months only means something if the gap is stable, and the interesting question is whether it is. If the frontier labs' next releases open it up again, then the open model is a lagging copy of a moving target. If it stays at three to six months across several cycles, then the target is being tracked, and the difference between the two is who gets to set prices.
What we would test
Two things, in order. First, reproduce the FLOPs and KV cache ratio against V3.2 on an independent serving stack at several context lengths, because that figure is the technical heart of the release and it is currently unverified outside the company. Second, run V4-Pro and the two closed models it is compared against on a task set the authors did not pick, and see whether three to six months survives contact with someone else's benchmark. The MIT licence means the first test costs only hardware. The second costs only honesty about what a scalar can tell you.
Sources
From the foundation