DashAttention and the year sparse attention became a field
DashAttention replaces fixed top-k block selection with alpha-entmax routing and matches full attention at 75 percent sparsity. Reading it alongside HISA and the DeepSeek indexer shows how quickly trainable sparse attention has split into distinct design camps.
Top-k was a budget, and budgets are the wrong prior
Every hierarchical sparse attention method of the last eighteen months shares one design. Chop the key-value cache into blocks, summarise each block cheaply, score the summaries against the query, keep the top k blocks, and run exact attention only inside them. NSA and InfLLMv2 both work this way. It is hardware friendly and it cuts the quadratic term, but k is a fixed budget, and the top-k operator is a hard cut with no gradient flowing back from the fine attention to the coarse routing decision.
The DashAttention paper from Tsinghua, Instituto Superior Técnico, CMU, Sapienza and Edinburgh, posted on 18 May, argues that both of those properties are mistakes. Some queries need many blocks and some need one, so a fixed k over-allocates easy queries and starves hard ones. And if the router cannot learn from the loss, the model never discovers which routing mistakes cost it accuracy.
What alpha-entmax changes
The replacement for top-k is alpha-entmax, a family of transformations that includes softmax at alpha equal to 1 and sparsemax at alpha equal to 2. For any alpha above 1, coordinates whose score falls below a data-dependent threshold become exactly zero. Sparsity is therefore a property of the score geometry rather than a hyperparameter. A sharp query with one obvious target routes to one block. A diffuse query routes to many.
DashAttention runs in three stages. Stage 0 builds a summary of each chunk using a small learned attention head initialised at zero, so training starts from mean pooling and only gradually learns something richer. Stage 1 applies alpha-entmax over the chunk scores to get a sparse routing distribution. Stage 2 runs a softmax over the tokens in the routed chunks, with the routing weights folded in as a prior on the logits. Because all three stages are differentiable, gradients from the final loss reach the summarisation and routing parameters. The authors use alpha of 1.5 at inference, increasing it from 1.25 during training.
The implementation is three fused Triton kernels, with AdaSplash-2 for the entmax stage and a FlashAttention style kernel for the token stage, so this is not a method that only exists in a proof-of-concept notebook.
The numbers
The experiments continue pretraining MiniCPM-4 base models at 1B, 3B and 8B on 16K context data and compare against full attention, NSA and InfLLMv2 at matched sparsity of 75 percent. On RULER at 16K, the 8B DashAttention model averages 83.6 against 85.3 for full attention, 55.0 for NSA and 78.9 for InfLLMv2. On HELMET the 8B numbers are 46.9 for DashAttention, 47.7 for full attention, 35.8 for NSA and 45.9 for InfLLMv2. Short-context general benchmarks are flat across methods, which is the expected result and still worth checking.
The gap widens with sparsity. On a HELMET sweep for the 8B model, DashAttention slightly exceeds full attention at low to moderate sparsity, and at roughly 90 percent sparsity it retains 39.4 percent overall accuracy, ahead of InfLLMv2 by 9 points and NSA by 19. On speed, prefill runs 1.34 to 3.09 times faster than FlashAttention-3 across 16K to 96K context, and decoding reaches 3.36 times faster at 96K with 93.75 percent sparsity.
One number we keep coming back to is in Table 4. A model trained with DashAttention and then served with plain softmax full attention scores higher than a model trained with full attention throughout, on both RULER and HELMET at every size. Sparse training appears to act as a regulariser on where the model looks, and the benefit survives even when the sparsity is removed at inference.
Non-dispersive attention at long context
The theoretical section explains why the routing choice matters more as context grows. Softmax attention is dispersive: as sequence length n grows, the entropy of the attention distribution grows like log n, so probability mass spreads thinner over more positions and long-range retrieval gets harder. Top-k bounds the entropy by log k, which helps, but the existing methods aggregate heads with softmax before the top-k cut, and that aggregation reintroduces dispersion.
The paper's Theorem 1 states the contrast. Softmax head aggregation is dispersive for any finite number of heads. Entmax head aggregation, provided each head's support grows sublinearly with n, is not. The practical consequence shows up in the RULER multi-key retrieval tasks, MK1 to MK3, where DashAttention's margins over the baselines are largest. There is also a nice figure showing per-layer sparsity on a RULER input: early layers stay dense and middle layers go sparse, an allocation the model found on its own that resembles hand-designed layer budgets from prior work.
The other camp: token-level indexing
DashAttention belongs to the block-routing camp. The other camp scores every token. DeepSeek Sparse Attention, used in DeepSeek-V3.2 and adopted by GLM-5, runs a lightweight indexer over every historical key for each query and forwards the top k tokens to a sparse latent attention operator. That gives fine-grained selection but the indexer itself scans the full prefix, so per-layer indexing cost is quadratic again. The HISA paper from Peking University, revised in April, measures this directly: at 64K context with 2048 selected tokens, the sparse attention costs about 1.6 ms while the indexer costs 5.6 ms. The bottleneck moved into the indexer.
HISA's answer is to put a block filter in front of the token indexer. Pool each block's indexing keys, score the pooled keys, keep the top m blocks, then run the original token indexer only inside those blocks. It requires no training and leaves the downstream sparse attention untouched. The indexer kernel gets 2.16 times faster at a 4:1 compression ratio and up to 3.75 times faster under a fixed 8K budget at 64K context. On LongBench with DeepSeek-V3.2 the average moves from 51.05 to 50.78, and on GLM-5 HISA actually edges ahead of the original indexer, 46.32 to 46.01. On needle-in-a-haystack out to 128K it stays close to the original, while a block-only baseline collapses when the needle sits mid-context. HISA also cites IndexCache, which cuts indexer cost by reusing indices across nearby layers.
So the field has sorted into designs that route blocks with a learned differentiable router and designs that select tokens with a cheap indexer and then make the indexer hierarchical. They are not mutually exclusive. HISA's block filter is a hard top-m, and DashAttention's whole argument is that a soft entmax filter is better. The experiment we would like to see is HISA's indexer with an entmax first stage in front of DSA, trained rather than plugged in. Neither paper's code is integrated into vLLM or SGLang yet, and DashAttention's authors list that as their main open item.
Sources
From the foundation