Speculative decoding becomes standard: DSpark and multi-token drafters
Google shipped multi-token drafters for every Gemma 4 size in May, and this week DeepSeek published DSpark, the draft-and-verify system that has been running under V4. Notes on why the method is lossless, and on what actually decides the speedup you get.
From inference-team trick to a checkbox in the download
Two releases in the past two months moved speculative decoding from something an inference team bolts on to something that ships with the weights. On May 5 Google published multi-token prediction drafters for the Gemma 4 family, covering the E2B, E4B, 26B mixture-of-experts and 31B dense models, under Apache 2.0, with support in LiteRT-LM, MLX, Transformers, vLLM, SGLang and Ollama. The headline claim is up to a 3x speedup with no change in output quality, and the worked examples are more modest, around 2.2x on Apple Silicon at batch sizes of four to eight.
DeepSeek's contribution is DSpark, the draft model that has been running under DeepSeek-V4 in production, with a paper posted on July 6 by Xin Cheng and 32 co-authors and a full training and evaluation codebase called DeepSpec. The repository ships draft checkpoints for Qwen3 4B, 8B and 14B and for Gemma 4 12B, trained on the open version of PerfectBlend, which is 1.3 million samples weighted toward math and code. When two labs release drafters for each other's open models, the method has stopped being a competitive secret.
Why draft-and-verify does not change the output
The idea dates to Leviathan, Kalman and Matias in late 2022. Decoding one token per forward pass leaves a large model memory-bandwidth bound: each step reads every weight from HBM to produce a single token. A small drafter proposes several tokens cheaply, and the target model then scores the whole proposed block in one forward pass, which costs about the same as scoring one token because the weights are read once either way. Wherever the target agrees with the draft, you keep the token. At the first disagreement you take the target's own token instead and throw away the rest of the draft.
The part people skip is why this is exact rather than approximate. The original paper uses a rejection sampling rule: a drafted token is accepted with probability equal to the ratio of the target's probability to the drafter's, capped at one, and on rejection you sample from a corrected residual distribution. The proof shows the accepted sequence has exactly the target model's distribution. The drafter can be as bad as you like and the outputs remain samples from the big model. A bad drafter only costs you speed. That is the property that made it safe to turn on by default, and it is why Google can claim zero degradation without running an eval.
What DSpark adds
The bottleneck in practice is how many drafted tokens survive verification, and the DSpark paper is about pushing that number up. Its drafter is semi-autoregressive: a parallel backbone with 128-token sliding-window attention proposes a block of up to five tokens, and a small sequential head, a low-rank Markov module with rank 256, models the dependencies between positions so that later tokens in the block do not decay into noise. The drafter shares the frozen embedding and output head with the target.
The second piece is a confidence head, a linear projection with a sigmoid, that estimates for each position the probability the token survives verification given that the earlier ones did. A scheduler then decides how long a block to verify by treating it as a throughput problem across all requests on the engine. Under moderate load the paper reports verification budgets of four to six tokens per request, against a baseline of two. Under heavy load it shortens them, because verifying long speculative blocks for every user wastes compute that would serve more users.
On Qwen3-4B the mean accepted length is 5.70 on math, 5.38 on code and 3.54 on chat, against 4.62, 4.16 and 2.26 for Eagle3. Averaged across benchmarks DSpark beats Eagle3 by 30.9, 26.7 and 30.0 percent at 4B, 8B and 14B, and DFlash by 16.3, 18.4 and 18.3 percent. In production the paper reports per-user generation speedups of 60 to 85 percent for V4-Flash and 57 to 78 percent for V4-Pro at matched throughput.
What decides the speedup you actually see
Three things set the number in a real deployment. The first is acceptance length, and the numbers above show it is domain-dependent. Chat text is less predictable than code or arithmetic, so the same drafter accepts fewer tokens there. The second is batch size. The whole method works because a decode step is memory bound with room to spare. As batch grows, verification stops being free, and Google's own examples describe the speedup improving with batch on A100 only up to a point. DSpark's scheduler exists because at high concurrency the right answer is often to speculate less.
The third is what fraction of the wall clock is decode at all. An agent loop that reads a large repository spends most of its tokens in prefill, which is already compute bound and gains nothing from a drafter. Speculation helps the long generated diffs and the reasoning traces, and helps least when the model is mostly re-reading context.
Where this leaves agent workloads
For our own agent runs we would expect the gain to sit between the code and chat numbers, with the largest effect on long uninterrupted generations and almost none on tool-heavy turns where each generation is a few dozen tokens. The honest way to find out is to measure accepted length on your own traces rather than on MT-Bench, and DeepSpec makes that a short experiment now that the training code is open.
The thing we want someone to test is whether a drafter trained on a lab's own agent transcripts, rather than on a generic instruction mix, moves accepted length on tool calls, which have a very regular structure a small model should predict well. If it does, the speedup for agents could exceed the paper's chat numbers by a wide margin, and drafters would become part of the fine-tuning recipe rather than a serving detail.
Sources
- Google, Multi-token prediction drafters for Gemma 4
- DeepSeek, DeepSpec repository (DSpark, DFlash, Eagle3 draft models)
- Cheng et al., DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation (arXiv 2607.05147)
- Leviathan, Kalman and Matias, Fast Inference from Transformers via Speculative Decoding (arXiv 2211.17192)
From the foundation