The numbers

Inception released Mercury 2 on February 24 and the pitch is speed with reasoning attached. The company reports 1,009 tokens per second on Nvidia Blackwell GPUs and an end-to-end latency of 1.7 seconds, against 14.4 seconds for Gemini 3 Flash and 23.4 seconds for Claude Haiku 4.5 with reasoning on in the same comparison. Pricing is 0.25 dollars per million input tokens and 0.75 per million output, which The Decoder works out as half of Gemini 3 Flash on input and a quarter on output, and about four times cheaper than Haiku 4.5.

The benchmark table is the part that would have looked impossible a year ago for a diffusion model. GPQA Diamond 74, AIME 91, LiveCodeBench 67, SciCode 38, IFBench 71, TAU 53. These are Inception's own numbers and we have not seen an independent run, but they place the model in the speed tier rather than below it, and the whole point of the release is that the model gets there while decoding in parallel. The context window is 128k, it supports tool use and JSON output, and it is served behind an OpenAI-compatible API.

How a diffusion language model decodes

The mechanism goes back to Inception's paper from June 2025 on the original Mercury Coder models. An autoregressive model produces one token per forward pass, left to right, and its speed is bounded by how fast it can run that pass. A diffusion language model starts from noise over a whole block of positions and refines all of them at once over a small number of steps, coarse to fine, with a learned denoiser trained across noise levels. The backbone is still a transformer, which the paper describes as a deliberate choice so that the usual inference optimisations carry over.

Inception's analogy for Mercury 2 is that it is less typewriter and more editor revising a full draft at once. In the 2025 paper the Coder models reached 1,109 tokens per second for the Mini variant and 737 for Small on H100s, with quality that the authors put alongside Claude 3.5 Haiku on coding tasks and ahead of Codestral 2501 on fill-in-the-middle completion. Those were code models with no reasoning mode. Mercury 2 is a general model with what Inception calls tunable reasoning, and that is the new claim.

Can parallel denoising carry a chain of thought

The reason this is a real question is that chain-of-thought works because each step conditions on the ones before it. Step five of a proof depends on step four having been written. Autoregression gives you that ordering for free. A denoiser that is refining the whole block simultaneously has to discover the dependency structure through its refinement steps, and it is not obvious in advance that a few steps of parallel refinement can reproduce a long serial argument.

The AIME and GPQA scores are the first evidence that it can, at least at the length of a competition maths solution. What we would want to know, and what none of the public material says, is how the reasoning is laid out. If the model denoises a block, then commits it and denoises the next block conditioned on the committed text, then it is autoregressive at the block level and the chain of thought is carried by the block ordering, with the speedup coming from within-block parallelism. If it is refining a very long span end to end, that would be a much more surprising result. The company's description of refining multiple text blocks at once suggests the first, and it would explain why the speed and the reasoning can coexist.

The other thing we would want is the number of denoising steps at the reported quality. Diffusion speed comes from doing fewer sequential passes than there are tokens, and reasoning quality tends to improve with more passes. The 1,009 tokens per second figure is a point on that curve, and the benchmark scores are presumably a point on the same curve. Whether they are the same point is the question the launch material does not answer.

Where autoregression still wins

Three places, on the current evidence. The first is scaling. Inception's own paper says the scaling properties of diffusion language models are less well understood than the autoregressive ones, and the models it compares against are the small fast ones, not the frontier. Nobody has shown a diffusion model at the scale where the strongest reasoning models live, and the honest reading of the Mercury 2 table is that it competes with the models built for speed.

The second is the ecosystem. The whole stack of inference tricks built for autoregressive decoding, from speculative decoding to prefix caching to the KV cache itself, assumes tokens arrive one at a time. Some of it transfers because the backbone is a transformer, as Inception argues, and some of it does not. A model that is five times faster per token but cannot use a shared prefix cache across a thousand agent trajectories may not be five times cheaper in production.

The third is streaming. Users and agent frameworks have been built around tokens appearing in order. A model that produces a block at a time changes what partial output looks like, and everything from the user interface to the tool-call parser has to handle text that gets revised before it is final. That is an engineering problem rather than a research one, but it is the kind that decides adoption.

What we would test

The experiment that would tell us the most is cheap. Take the AIME problems, sweep the reasoning setting from lowest to highest, and record accuracy and tokens per second at each point. If accuracy climbs while throughput falls toward autoregressive rates, the reasoning is being bought with sequential steps and the headline numbers describe two different configurations. If accuracy holds at high throughput, parallel refinement is doing something that deserves a paper.

Either way, Mercury 2 is the first diffusion language model that we would consider putting into an agent loop, and that was not true of anything in this family a year ago. The speed-tier models it undercuts on price are the ones that do most of the volume in deployed systems. If the quality claims hold up independently, the interesting question stops being whether diffusion can reason and becomes how much of the token budget of the world it can take.

Sources

  1. The Decoder, Inception launches Mercury 2, the first diffusion-based language reasoning model
  2. Inception Labs, Mercury: Ultra-Fast Language Models Based on Diffusion (arXiv 2506.17298)
  3. Inception Labs, Introducing Mercury 2