mHC: constraining the residual stream so deep mixing stays stable
DeepSeek's New Year's Eve paper widens the residual stream, lets layers mix it, and then projects the mixing matrix onto doubly stochastic matrices so norms survive depth. A method explainer on why the residual connection is being redesigned and what it cost.
Why anyone would touch the residual connection
The residual connection is the one part of the transformer nobody argues about. Each layer adds its output to a running stream, and the stream is passed forward unchanged. That identity path is why gradients survive a hundred layers. Hyper-connections, an earlier idea the DeepSeek paper builds on, ask whether unchanged is too conservative. They widen the stream by a factor n, typically 4, so the model carries n parallel copies of the hidden state, and they let the model learn how to read from, write to, and mix those copies.
Three learned maps do this. H_pre, a 1 by n vector, aggregates the n copies into a single input for the layer. H_post, another 1 by n vector, spreads the layer's output back across the copies. H_res, an n by n matrix, mixes the copies with each other in the stream itself. The paper reports that H_res is where most of the gain comes from. It also does not change per layer FLOPs, since the attention and MLP blocks still see a C dimensional input.
What goes wrong without a constraint
The problem is composition. The stream at layer L has been multiplied by H_res at every layer before it, so the effective map is a product of L learned matrices. Nothing stops that product from growing or shrinking. The authors measured this on unconstrained hyper-connections and found the maximum gain magnitude of the composite map peaking around 3000, where the ideal value is 1. A signal amplified by that much somewhere in the stack, or attenuated by its inverse, is what training instability looks like from the inside.
This is the identity mapping property that the plain residual gets for free. With a single copy and no mixing, the product of L identities is the identity, and the gradient path is clean. Widening the stream broke that guarantee and the fix has to restore it without giving up the mixing.
The manifold
The mHC answer is to constrain H_res to be doubly stochastic. Every entry non negative, every row summing to 1, every column summing to 1. That set is the Birkhoff polytope, and it has two properties that matter here. Its spectral norm is bounded by 1, so applying it cannot blow up the signal. And it is closed under multiplication, so a product of doubly stochastic matrices is still doubly stochastic. The composite map across a thousand layers stays inside the same well behaved set.
The geometric picture is that a doubly stochastic matrix is a convex combination of permutations. Mixing the n copies of the stream is allowed, but only as a weighted average of ways to shuffle them, which conserves the total signal. In practice the model computes an unconstrained matrix and runs 20 iterations of Sinkhorn-Knopp, alternately normalising rows and columns of the exponentiated matrix, to project it onto the polytope. H_pre passes through a sigmoid and H_post through 2 times a sigmoid so both stay non negative.
Measured on the same composite map, the maximum gain magnitude drops from about 3000 to about 1.6. Gradient norm curves for mHC track the baseline model rather than the spiky hyper-connection run, and the paper says the instability observed in HC is effectively mitigated.
What it cost and what it bought
The overhead is memory traffic rather than compute, since the stream is now n times wider and every read and write to it is n times larger. The paper's infrastructure section is as long as the method section for that reason. They fuse the mixing operations into custom TileLang kernels, move the RMSNorm after the matrix multiply, and recompute the expanded states during backpropagation rather than storing them, with a block size chosen to line up with pipeline stage boundaries. With n equal to 4 the reported time overhead in large scale training is 6.7 percent.
The models are mixture of experts in the DeepSeek-V3 style, at 3B, 9B and 27B total parameters, with 612M, 1.66B and 4.14B active, sequence length 4096, and data scaled with model size, about 262B tokens for the 27B run. On that model the baseline, HC and mHC score 43.8, 48.9 and 51.0 on BBH, and 46.7, 53.2 and 53.8 on GSM8K. The gain over plain HC is modest, around 2 points on BBH and DROP. The gain over the baseline is larger, and it now comes with a run that does not diverge.
Reading the tea leaves
Sebastian Raschka's timeline of DeepSeek's technical releases places this paper on December 31, 2025, after the V3.2 line that introduced sparse attention in September and the refined V3.2 training on December 1. He describes mHC as recent architecture research and does not tie it to any announced model, and we will not either. A paper that spends this much effort on pipeline scheduling and kernel fusion is a paper from a lab that intends to train with it, and that is as far as the evidence goes.
The result we would want to see is depth. The stability argument is about products of many matrices, and the paper's experiments do not appear to push layer count beyond ordinary values. If the constraint really preserves the identity property across a thousand layers, a small model trained very deep with mHC against very deep HC would show it directly. Someone with a modest cluster could run that this month.
Sources
From the foundation