Three bugs and a batch: the Anthropic postmortem and the nondeterminism paper
Anthropic's postmortem traces weeks of quality complaints to a routing error, a TPU runtime bug and a mixed precision top-k miscompilation. A week earlier Thinking Machines showed that batch size, not floating point noise, is why the same prompt gives different answers. Both say quality drifts in serving, and the weights never changed.
The weights were fine
For most of August and into September, people using Claude reported that it had got worse, and Anthropic's answers in that period did not satisfy anyone. On September 17 the company published a postmortem, and the first thing it says is that the models themselves never changed. Three separate infrastructure bugs, overlapping in time, on different hardware, produced different symptoms on different platforms at different rates. Reading it alongside the Thinking Machines piece from a week earlier, we think the pair form the clearest account yet of a fact the field has been slow to accept. Model quality lives in the serving stack as much as in the checkpoint.
Bug one: the wrong server
The first bug was a routing error introduced on August 5. Some short context requests for Sonnet 4 were sent to servers configured for the one million token context window. At first it touched 0.8 percent of Sonnet 4 requests. By August 31 it had peaked at 16 percent. Anthropic estimates that about 30 percent of Claude Code users who made requests in the window had at least one message routed to the wrong server type. Bedrock saw a peak of 0.18 percent and Vertex under 0.0004 percent, which is why complaints clustered on first party users and why cross platform comparisons were so confusing. The fix went in on September 4 and finished rolling out by September 18.
Bugs two and three: the TPU path
The second bug was a runtime performance optimisation deployed to TPU servers on August 25 that corrupted token generation. The visible symptom was Thai or Chinese characters appearing in English answers, and syntax errors in code. It affected Opus 4.1, Opus 4 and Sonnet 4 on first party serving only and was rolled back on September 2.
The third is the one worth studying. On the same day, August 25, a change on the TPU serving path exposed a latent bug in the XLA:TPU compiler. The compiler's approximate top-k operation was mixing bf16 and fp32 arithmetic in a way that sometimes changed which token was selected, and because the miscompilation depended on the surrounding code, it came and went. The optimisation pass responsible is guarded by a flag called xla_allow_excess_precision, which defaults to true. Haiku 3.5 was rolled back on September 4, Opus 3 on September 12, and Sonnet 4 shortly after out of caution. A precision decision buried in a compiler default changed which token a model picked, on some requests, for some batch shapes.
Why nobody could see it
The section on detection is the most honest part of the document. Anthropic's evaluations did not capture the degradation, partly because Claude often recovers well from isolated mistakes, so a single dropped token rarely fails a benchmark. Privacy controls limit how much engineers can look at actual user conversations, so the signal that users had was not available to the people debugging. And the three bugs had different symptoms on different platforms, so any one report contradicted another. The changes announced are what you would expect. Evaluations designed to separate working from broken implementations rather than to rank models, continuous evaluation on production systems rather than periodic checks, and better tooling for turning community reports into something an engineer can act on without reading private chats.
The batch was the variable all along
Horace He's piece for Thinking Machines, published September 10, explains a related mystery that has been treated as a law of nature. Ask a model the same question at temperature zero twice and you often get two answers. The usual story blames floating point non-associativity plus GPU concurrency. He shows that story is incomplete, because a matrix multiply run repeatedly on the same GPU gives bitwise identical results. The real cause is that inference kernels are not batch invariant. The reduction order inside RMSNorm, matrix multiplication and attention changes with batch size, and an inference server's batch size depends on who else is sending requests at that moment. Your answer depends on the other users' traffic.
The demonstration is clean. Sampling 1,000 completions from Qwen3-235B at temperature zero produced 80 distinct outputs. All 1,000 agreed on the first 102 tokens, then split at token 103, where most wrote Queens, New York and eight wrote New York City. With batch invariant versions of the three kernels, all 1,000 completions were identical. The cost was real but survivable, 55 seconds against 26 for stock vLLM in a naive version and 42 seconds with a better attention kernel. As a bonus, deterministic inference lets reinforcement learning be exactly on policy, with zero KL divergence between the sampling and training policies, which off policy setups can only approximate.
What the two pieces say together
Put them side by side and the message is that the last mile of serving is a place where numerics silently decide outputs. Anthropic's third bug was batch dependent precision loss inside a compiler. He's nondeterminism is batch dependent reduction order inside kernels. Neither is visible in the weights, neither is caught by a standard benchmark, and both change which token comes out. The difference is that one was an accident and the other is the default state of every inference server in the world.
What we would want from labs now is a serving eval that runs the same fixed prompt set continuously against production, at varying batch sizes, and alarms on any distribution shift in the outputs, because both documents show that is where the drift hides. And we would want the batch invariant kernels to become a standard option in vLLM and SGLang, so that anyone running an evaluation can turn them on and know that the number they publish is a property of the model and not of whoever else was using the GPU that afternoon.
Sources
From the foundation