The month the transformer got competition: Mamba and Mixtral
Two December releases attack the transformer from different sides. Mamba replaces attention with a selective state space, and Mixtral keeps attention but routes each token through two of eight experts.
Two papers, ten days apart
Albert Gu and Tri Dao posted Mamba on December 1. Mistral released Mixtral 8x7B on December 11. Both are pitched as ways to get transformer-class quality at a fraction of the compute per token, and they get there by removing different things. Mamba removes attention. Mixtral removes the idea that every parameter should touch every token.
We have spent the last week reading both and running the released Mamba checkpoints. What follows is our attempt to sort out what each one has actually demonstrated, as opposed to what people on social media are saying it demonstrates.
What Mamba changes
State space models had a known problem on text. Earlier SSMs like S4 used time-invariant parameters, so the recurrence treated every token the same way regardless of content. That is fine for audio and poor for language, where the model needs to decide token by token what to keep and what to drop. Gu and Dao's fix is to make the SSM parameters, specifically the step size and the input and output projections, functions of the current input. They call this selection. In their words it lets the model selectively propagate or forget information along the sequence length dimension depending on the current token.
Making the parameters input-dependent breaks the convolutional trick that made S4 fast to train. Their answer is a hardware-aware parallel scan. They load the SSM parameters from HBM into SRAM, run the discretization and recurrence there, and write only the outputs back. Kernel fusion, parallel scan, and recomputation in the backward pass keep the expanded state out of GPU memory. The result is a recurrent model that trains in linear time and, because there is no KV cache at inference, can use much larger batches. They report 4 to 5 times the inference throughput of a comparable transformer.
The synthetic tasks are the cleanest evidence that selection works. On selective copying the model has to remember specific tokens and ignore filler, which time-invariant SSMs cannot do. On the induction heads task, Mamba trained on short sequences generalizes to million-length sequences, 4000 times longer than anything seen in training, while the paper reports no other method going beyond 2 times.
The language modeling numbers
On the Pile, Mamba-1.4B reaches 6.80 perplexity against 7.51 for Pythia-1.4B, and Mamba-2.8B reaches 6.22 against 6.73 for Pythia-2.8B. On downstream evaluations the authors say Mamba-3B matches transformers twice its size. These are real gains at the scales tested. The largest model in the paper is under 3B parameters, and the comparison baseline is Pythia, which is a 2023 open replication rather than a tuned frontier transformer.
So the honest statement is that at small scale, with matched data, a pure selective SSM beats a transformer of the same size and there is no attention anywhere in the network. Whether that gap holds, closes, or reverses at 30B or 70B is unknown and is the single most important open question the paper leaves. The authors are explicit that there is no free lunch across modalities either. They note that complex-valued state helps continuous data like audio and hurts discrete data like text and DNA, so the architecture that wins on one may not win on the other.
What Mixtral changes
Mixtral 8x7B is a decoder-only transformer where each feed-forward block is replaced by eight experts and a router picks two of them per token. The full model has 46.7B parameters but only 12.9B are used for any given token. Mistral says it outperforms Llama 2 70B on most benchmarks with 6 times faster inference, and matches or beats GPT-3.5 on most standard benchmarks. The instruction-tuned version scores 8.30 on MT-Bench. It handles 32k context, works in English, French, Italian, German and Spanish, and ships under Apache 2.0.
Sparse mixture of experts is not new, but an open-weight MoE at this quality with a permissive license is. The trade is memory for compute. You still have to hold all 46.7B parameters somewhere, which is why it does not run on the hardware a 13B dense model runs on, but each forward pass costs roughly what a 13B model costs. For anyone serving at volume that is the right trade, and for anyone finetuning on a single consumer GPU it is the wrong one.
The announcement is a product release rather than a paper, so there is a lot we do not know. There is no public training data description, no ablation on the number of active experts, and no analysis of what the router learned. We would treat the benchmark table as a claim to be verified rather than a result, and we expect independent evaluations within weeks.
How we read the two together
Mixtral is the safer bet. It keeps the transformer and every piece of tooling built around it, and it makes a well-understood technique available in the open. Mamba is the riskier and more interesting one, because if selective SSMs keep pace with attention at scale, the quadratic cost of long context stops being a law of nature. Our guess is that the near future is hybrids, with a few attention layers for precise retrieval and SSM layers for everything else, because the induction task shows that selection can do something attention does badly, and the small-scale comparisons do not yet show it can do everything attention does well.
The experiment we would most like to see from either camp is a matched 7B run, same data, same tokens, Mamba versus a dense transformer versus a Mixtral-style MoE, with long-context evaluations that go past the training length. That would settle more than any amount of argument about which December release mattered more.
Sources
From the foundation