Gemini Diffusion: text generation without the left-to-right rule
Google's we/O research demo generates code by denoising whole blocks at 1,479 tokens per second and lands within a point or two of Gemini 2.0 Flash-Lite on most code benchmarks. An explainer on how diffusion applies to text, what it gives up, and why it matters.
The demo
At we/O on May 20, Google DeepMind showed an experimental model called Gemini Diffusion that produces text and code by refining noise into output instead of predicting one token after another. The model page lists a sampling speed of 1,479 tokens per second with 0.84 seconds of overhead latency. Simon Willison got access, asked it to build a simulated chat app, and reports the whole interactive HTML and JavaScript page arriving in single-digit seconds at 857 tokens per second in his run. His summary was that they are not kidding about it being fast.
The speed is the reason the demo exists, but the benchmark table is what makes it interesting. Against Gemini 2.0 Flash-Lite, Google's fastest production model, the diffusion model scores 30.9 versus 28.5 on LiveCodeBench v6, 45.4 versus 45.8 on BigCodeBench, 56.8 versus 56.0 on LBPP v2, 89.6 versus 90.2 on HumanEval and 76.0 versus 75.8 on MBPP. On coding tasks the two are within a point of each other, and the diffusion model gets there in a fraction of the time. Access is by waitlist.
How diffusion applies to text
Image diffusion works by adding Gaussian noise to a picture until it is static, training a network to reverse one step of that process, and then running the reversal from pure noise. Text is discrete, so you cannot add a little Gaussian noise to a token. The version that works for language, and the one Willison points to, is much closer to BERT's masked language modelling. The corruption process replaces tokens with a mask, up to and including masking everything. The model is trained to predict the masked tokens given the unmasked ones. Generation starts from a fully masked block of some fixed length and fills it in over a series of steps, each step committing to some tokens and leaving others for later.
Two properties follow. First, every step predicts many positions in parallel, so the number of forward passes is the number of denoising steps rather than the number of tokens, and that is where the throughput comes from. Second, tokens that were filled in early can be revised in later steps. Google's page describes this as iterative refinement that corrects errors during generation. An autoregressive model has no such mechanism. Once it emits a token, every later token is conditioned on it.
What it gives up
The same benchmark table shows where the tradeoff lands. On GPQA Diamond the diffusion model scores 40.4 against 56.5 for Flash-Lite. On BIG-Bench Extra Hard it is 15.0 against 21.0. On Global MMLU Lite it is 69.1 against 79.0. On SWE-Bench Verified, the one agentic coding benchmark in the table, it is 22.9 against 28.5. AIME 2025 goes the other way, 23.3 against 20.0, but the pattern across knowledge and multi-step reasoning tasks is a clear deficit. Code with a fixed structure suits parallel infilling. Long chains of dependent reasoning may not.
There are structural costs the table does not show. A block of fixed length has to be chosen before generation starts, so the model has to know, or guess, roughly how long its answer is, and generating past the block means another block. The KV cache that makes autoregressive decoding cheap per token does not apply in the same way, because the context is re-encoded each step. And the well-tuned inference stack for autoregressive models, speculative decoding, prefix caching, streaming token by token, was built for autoregressive decoding. Some of that will transfer and some will have to be reinvented.
Why this one matters
Text diffusion is not new as a research direction and Google is not the first to ship one. Willison notes that Inception Labs released a commercial diffusion model called Mercury in February. What is new is that a frontier lab has put a diffusion model on the same benchmark table as one of its own production autoregressive models and shown parity on a real task family. That changes the question from whether diffusion can produce coherent text to which workloads justify the tradeoff, and that is a question engineering teams can answer.
For a foundation like ours the interesting part is the training objective. If masked prediction over whole blocks reaches autoregressive quality on code at a fraction of the decode cost, then the left-to-right factorisation that every large model since GPT-2 has used was a choice, and it can be revisited. We do not expect it to be replaced. We expect the two to coexist, with diffusion handling latency-sensitive generation of structured output and autoregression keeping the long reasoning traces, and we expect hybrids where a diffusion model drafts and an autoregressive model verifies.
What we would want to see next
The obvious missing number is the cost. Tokens per second is a throughput figure and tells you nothing about compute per token, which depends on the number of denoising steps and the block length, neither of which Google has published. A model that is five times faster to the user and three times more expensive to serve is a different product from one that is faster and cheaper. We would also want to see how quality changes as the step count is reduced, since that curve is where the practical speed lives.
The experiment we would run is on revision. The claim that later steps correct earlier mistakes is the one property autoregressive models cannot have, and it should be measurable. Take code outputs, log which tokens changed between steps, and check whether the changes fix bugs or introduce them. If the edits are real repairs, that is a stronger argument for diffusion than any throughput number, because it means the model is doing something in generation that a left-to-right decoder structurally cannot.
Sources
From the foundation