DeepSeek V3.2 and the lightning indexer
DeepSeek V3.2 makes sparse attention the default. A small FP8 indexer scores past tokens and the main attention only touches the top 2,048. Notes on how the mechanism works, how it was trained in, and what the Speciale variant proves.
From experiment to default in two months
DeepSeek released V3.2-Exp on September 29 as an experimental build on top of V3.1-Terminus, with a new attention mechanism called DeepSeek Sparse Attention and an API price cut of more than 50 percent effective the same day. They kept V3.1-Terminus reachable on a temporary endpoint until October 15 so people could run their own comparisons, and their own benchmarks put the two on par. The kernels were released in TileLang and CUDA.
On December 2 the full V3.2 paper landed on arXiv with over 260 authors, and sparse attention is no longer an experiment. It is the architecture. The paper describes three things: the sparse attention mechanism, a scaled up reinforcement learning stage, and a synthesis pipeline for agentic tool use data. This post is about the first one, because it is the part that changes the cost of a long context and the part most likely to be copied.
How the lightning indexer works
Standard attention compares each query to every preceding token, which is the quadratic term everyone has spent years trying to remove. DSA keeps that comparison but does it twice, once cheaply and once properly. The cheap pass is the lightning indexer, a small set of indexer heads that compute a relevance score between the query and each earlier token as a weighted sum of ReLU activated dot products. It runs in FP8 and the paper attributes its efficiency to that choice.
The indexer then selects the top 2,048 key value entries for each query, and the main attention runs only over those. That takes the main term from order L squared to order L times k with k fixed at 2,048. The indexer itself is still quadratic in sequence length, and the authors do not pretend otherwise. Their claim is that its cost is far below the cost of the multi-head latent attention it is gating, so the overall budget drops even though one quadratic term survives.
One detail we found clarifying is that DSA sits on the MQA mode of MLA. Each latent vector, which is the key value entry in MLA, is shared across all query heads. So the indexer selects latents, not per-head keys, and the selection is made once per query position rather than once per head. That is what keeps the bookkeeping cheap enough to be worth doing.
Teaching a dense model to be sparse
The training recipe is continued pretraining from a dense checkpoint in two stages. In the dense warm-up, every parameter is frozen except the indexer, and the indexer is trained with a KL divergence loss to match the distribution of the main attention. That runs 1,000 steps at learning rate 1e-3, about 2.1 billion tokens. The indexer learns to predict where dense attention would have looked before it is ever allowed to restrict it.
Then the sparse stage turns on top-k selection and unfreezes the whole model at a learning rate of 7.3e-6 for 15,000 steps, about 943.7 billion tokens. The total is roughly 946 billion tokens to convert a dense model into a sparse one, which is a lot of tokens and a small fraction of a pretraining run. The lesson we take is that sparsity was learned as a compression of an already trained attention pattern rather than discovered from scratch, and that ordering is probably why it did not cost quality.
What the numbers say
The paper's comparison table puts V3.2 at 93.1 percent on AIME 2025 against 94.6 for GPT-5 and 95.0 for Gemini 3.0 Pro, at 92.5 on HMMT February 2025 against 88.3 and 97.5, and at 73.1 on SWE-bench Verified against 74.9 and 76.2. Those are close enough that sparse attention is not visibly costing anything on the tasks people care about, and the paper's headline claim is that the reinforcement learning stage, run with GRPO at a post-training budget exceeding 10 percent of pretraining cost, gets the model to GPT-5 level.
V3.2-Speciale is the variant that spends more at inference. It scores 96.0 on AIME and 99.2 on HMMT, and the paper reports gold medal level results on the 2025 International Mathematical Olympiad and the International Olympiad in Informatics. We read Speciale as an existence proof rather than a product. It shows the sparse architecture has headroom when you let it think longer, and it puts an open weights model in territory that until recently belonged to closed labs.
What we would test
The obvious experiment is whether k equals 2,048 is a property of this model or a property of the task. A fixed top-k is a fixed information budget per query, and we would expect tasks that need diffuse attention over a long document, like summarising a hundred page contract, to suffer before tasks that need a few sharp lookups. The paper says the mechanism preserves performance on long context, and the benchmark table backs that up for reasoning, but we have not seen a breakdown by attention pattern.
The other one is the indexer's own quadratic term. At 2,048 selected tokens and a million token context, the indexer is doing all the work and the main attention is nearly free. Whether the FP8 indexer stays cheap at that length, or whether it becomes the thing you next need to make sparse, is the question that decides how far this design carries. DeepSeek has released the kernels, so someone will measure it soon.
Sources
From the foundation