The claim in the technical report

The Kimi K2 report, posted to arXiv on July 28 with 199 authors, makes a claim that anyone who has babysat a large pretraining run will read twice. The model has 1.04 trillion total parameters, 32.6 billion of them active per token, and it was pretrained on 15.5 trillion tokens. The team says the training loss stayed smooth for the whole run, with no observable spikes. They attribute that to an optimizer they call MuonClip, which is Muon plus a weight rescaling trick they name QK-clip.

The architecture is close to DeepSeek-V3 in spirit. There are 61 layers with a hidden dimension of 7168, 384 experts of which 8 are active for any token, which the team describes as a sparsity ratio of 48, and multi-head latent attention with 64 heads. The head count is half of what DeepSeek-V3 uses, and the report says that choice was made for inference efficiency. None of that is the interesting part. The interesting part is that the optimizer story is the headline, and that a Chinese lab chose to tell it in full.

What Muon does differently from AdamW

AdamW keeps a running estimate of the first and second moments of each parameter's gradient and divides one by the square root of the other. It treats every parameter as an independent scalar. Muon treats a weight matrix as a matrix. It takes the momentum-accumulated gradient for a layer and replaces it with the nearest orthogonal matrix, computed approximately with a few Newton-Schulz iterations, so that every direction in the update has roughly equal magnitude regardless of how strongly the gradient favoured it.

The K2 report writes the update as the Newton-Schulz output multiplied by 0.2 times the square root of the larger matrix dimension, which is a scaling chosen to match the root mean square size of an Adam update. That detail matters in practice because it means Muon can be dropped in with learning rates and weight decay that were tuned for Adam. The report states that Muon reaches a given loss in fewer tokens than AdamW under the same compute budget, which is the token efficiency claim. It does not give a single headline percentage for K2, so we will not invent one.

The intuition for why orthogonalized updates help is that a raw gradient for a large matrix is dominated by a few large singular directions, and the many small directions barely move. Orthogonalizing flattens the spectrum so the small directions get updated too. Whether that is the whole explanation is an open question, and the K2 report does not claim to have settled it.

Why the attention logits explode

The cost of that flattened spectrum shows up in attention. The report describes a failure mode where the dot products between queries and keys grow without bound as training proceeds, until the softmax saturates and the model stops learning through those heads. This happens with vanilla Muon at scale, and the team traces it to the way orthogonalized updates keep pushing weight norms outward in every direction.

The standard fixes do not fit their architecture. Logit soft-capping bounds the logits but does not stop the underlying weights from growing. Query-key normalization, which normalizes the query and key vectors before the dot product, assumes you can get at the full key matrix, and in multi-head latent attention the keys are never fully materialized because they are reconstructed from a compressed latent. So the team needed something that acts on the weights rather than the activations.

They ran a mid-scale test to show the problem is real. On a mixture of experts with 9 billion active and 53 billion total parameters, vanilla Muon produced maximum attention logits that rapidly exceeded a magnitude of 1000, and training became unstable. That is the experiment we would want anyone to reproduce before trusting the rest of the story.

What QK-clip actually does

QK-clip is a rescaling applied to the projection weights after each optimizer step. For each attention head the team tracks the maximum logit seen in that step. If it exceeds a threshold tau, set to 100 for K2, the head's query and key projections are scaled down by a factor gamma equal to tau divided by the observed maximum, capped at 1 so that well-behaved heads are left alone. The clip is per head, so one runaway head does not drag down the others.

The MLA-specific detail is where the care shows. In MLA the query and key each have a head-specific component and a shared rotary component. The team applies the clip only to the head-specific parts, scaling each by the square root of gamma so the product comes out at gamma, and leaves the shared rotary component untouched. That avoids the case where clipping one head's logits silently changes the position encoding that every other head depends on.

With this in place, the report says the full K2 run showed no spikes at all. What we like about the mechanism is that it is cheap, it only fires when needed, and it leaves no trace once training is done. What we would want to know is how often it fired, and whether the same heads kept tripping it. The report does not say.

On who told the story first

Muon was not invented at Moonshot. It came out of the open speedrun community and had been discussed for months as a small-model curiosity. What Moonshot did was run it at a scale nobody else had published, hit the instability nobody had documented at that scale, and write down the fix in enough detail to reimplement. The Western frontier labs have optimizers and stability tricks of their own, and we do not know what they are, because they are not in any report.

That is the part we keep coming back to. The optimizer is one of the last places where a lab's private engineering knowledge still translates into a compute advantage, and a lab that wants to be taken seriously as an open contributor just gave that knowledge away. If the token efficiency claim holds up in independent runs, the next few open pretraining efforts will all use some version of this, and they will use it because of a technical report from Beijing. We would want to see someone train a matched pair, AdamW and MuonClip, at a few billion parameters with the same data and publish both curves. Until then the claim is one lab's word, and a very specific one.

Sources

  1. Kimi Team, Kimi K2: Open Agentic Intelligence (arXiv 2507.20534)