DeepSeek-V3: the 2.8 million GPU-hour frontier model
A 671B mixture-of-experts model trained on 14.8 trillion tokens for a reported 2.788 million H800 GPU hours. A look at which engineering choices in the technical report actually drove that number, and which are just good ideas that happened to ship in the same model.
The number and what it covers
The technical report DeepSeek posted on December 27 gives a training cost table that most frontier labs do not publish. Pre-training took 2,664 thousand H800 GPU hours. Context extension, in two stages to 32K and then 128K, took 119 thousand. Post-training took 5 thousand. The total is 2.788 million GPU hours, and at an assumed rental price of 2 dollars per H800 hour that is 5.576 million dollars. Per trillion tokens of pre-training the figure is 180 thousand GPU hours, or 3.7 days on the 2,048-GPU cluster they used.
The report is careful about what the number excludes. It covers the final training run only, and leaves out the cost of prior research and ablation experiments on architectures, algorithms and data. So this is the price of the run. The price of the model, research included, is higher and unreported. The run figure is still a much more useful disclosure than nothing, and it is the number to hold in mind when reading the rest of this piece.
The model itself is 671 billion parameters with 37 billion active per token, trained on 14.8 trillion tokens, with no irrecoverable loss spikes and no rollbacks reported over the whole run. The architecture reuses Multi-head Latent Attention and DeepSeekMoE from V2, and adds two new pieces, an auxiliary-loss-free load balancing strategy and a multi-token prediction objective. The systems side adds FP8 mixed-precision training and a pipeline scheduler called DualPipe.
Sparsity did most of the work
It helps to separate the contributions. The single biggest lever is that only 37 billion of the 671 billion parameters run on each token. Compute per token scales with active parameters, so the run is priced like a model a fraction of the total size. That is the mixture-of-experts bet, and it predates V3. Everything else in the report is about making that bet pay off without the usual costs of MoE training.
The usual costs are two. Experts collapse, with a few experts taking all the traffic while others idle, and the standard fix is an auxiliary loss that penalises imbalance and in doing so hurts the model. Communication across nodes for expert routing eats into the compute you saved. V3's contributions map onto those two problems.
Auxiliary-loss-free balancing
Instead of a balance loss, V3 keeps a bias term per expert that is added to the routing score only when choosing which experts to activate. The bias is not used in the gating value that multiplies the expert output. After each step, an expert that was overloaded has its bias decreased by a fixed amount and an underloaded expert has it increased. Balance is enforced by nudging the routing, and the gradient the model sees is untouched. A small sequence-wise balance loss is kept as a backstop for extreme imbalance within a single sequence.
The ablation in the report compares this to the auxiliary-loss approach and finds better performance at the same balance. What it buys in cost terms is that the model does not spend capacity satisfying a loss term unrelated to language modelling. We would file this as a quality improvement at fixed compute rather than a compute saving, but it is what makes the aggressive sparsity usable.
FP8 and the pipeline
The direct compute saving is FP8. The report describes a mixed-precision framework in which most matrix multiplications run in 8-bit floating point, with activations quantised on 1 by 128 tiles and weights on 128 by 128 blocks so that outliers in one channel do not destroy the scale for the whole tensor. Activations are cached and dispatched in FP8, and low-precision optimiser states are stored in BF16. They validated the scheme at two smaller scales for about a trillion tokens each and report a relative loss error against a BF16 baseline that stayed below 0.25 percent.
Sub-quarter-percent loss error for a roughly halved memory and bandwidth footprint on the heaviest operations is the kind of trade that shows up directly in GPU hours. This is the first frontier-scale run we know of that reports training the main model in FP8 rather than only using it for inference.
DualPipe attacks the communication problem. It overlaps computation and communication within pairs of forward and backward chunks and schedules micro-batches from both ends of the pipeline, which the report says leaves fewer bubbles than the standard 1F1B schedule at the cost of holding two copies of the parameters. Combined with custom cross-node all-to-all kernels, the aim is that expert routing traffic across InfiniBand never stalls the GPUs. Without this, the sparsity saving would leak back out as idle time.
Multi-token prediction is a quality trick
The multi-token prediction objective adds small sequential modules that predict the second next token from the main model's hidden state, sharing the embedding and output head. It densifies the training signal and the report's ablation shows it helps benchmark scores. At inference the modules can be dropped or reused for speculative decoding. It does not reduce the cost of the run, and if anything adds a little. It is in the paper because it makes the model better.
So our accounting of the 2.788 million hours is roughly this. Sparsity sets the order of magnitude. FP8 takes a large constant factor off the top. DualPipe and the communication kernels stop the sparsity saving from being eaten by network stalls. Auxiliary-loss-free balancing and MTP are what let a model trained that cheaply be good enough to compare with closed models. What we would want someone to try is a re-run of the FP8 ablation on a different architecture, because if the 0.25 percent figure holds generally, FP8 pre-training should be the default everywhere within a year.
Sources
From the foundation