The shape of the model

AI21 released Jamba on March 28 with a technical report and Apache 2.0 weights. It has 52 billion total parameters and 12 billion active, and it is built from repeating blocks of eight layers. In each block, one layer is standard attention and seven are Mamba layers, the selective state space design from December. Every other layer carries a mixture-of-experts feed-forward with 16 experts and top-2 routing. That is the whole recipe: one attention layer for every seven Mamba layers, MoE on half of them.

The headline consequence is memory. Attention keeps a key-value cache that grows with context length, and at 256K tokens that cache is what makes a Mixtral or Llama-2 70B deployment expensive. With only one attention layer in eight, Jamba's cache is a fraction of the size, which is how a 52B model with a 256K trained context fits 140K tokens of context on a single 80GB GPU. The report puts throughput at long context at about three times Mixtral. The active parameter count is in the same range as Mixtral's, so the difference comes from the cache rather than from the compute per token.

What the ablations found

The report is more useful for its ablations than for its benchmark table, and three findings stand out. The first is the ratio. AI21 compared one attention layer per three Mamba layers against one per seven and found the 1:3 ratio scored better. They shipped 1:7 anyway because the memory and throughput gains at long context were worth the small quality cost. That is a deliberate engineering trade rather than the best point on the quality curve, and the report says so.

The second is that pure Mamba, with no attention at all, had a specific failure. It could not do in-context learning in the way a transformer can. On tasks like IMDB sentiment, QuAC and NarrativeQA it would often produce an answer in the wrong format, ignoring the pattern set by the few-shot examples. Add a small number of attention layers and the behaviour returns. The authors' reading is that attention provides the copying and induction machinery that in-context learning relies on, and a state space layer does not readily learn it.

The third is positional encoding. Jamba uses no explicit positional embeddings, no RoPE, and the ablation showed adding them did not help. The Mamba layers carry position implicitly through their recurrence, so the attention layers get what they need from the surrounding context. The one stability fix they needed was RMSNorm inside the Mamba layers, which stopped loss spikes at scale.

Why hybrid rather than pure

When Mamba came out the interesting question was whether state space models would replace attention outright. Jamba is the first large-scale answer, and the answer is no, or at least not yet. The in-context learning gap is the reason. A model that handles long documents cheaply but cannot follow a few-shot format is not a usable language model, and the cheapest fix is one attention layer in eight.

That fix is a good deal for the attention side too. Most of the quadratic cost goes away, the KV cache shrinks by nearly an order of magnitude, and what remains is enough for the retrieval and copying behaviour that makes transformers useful. We expect this ratio to be tuned rather than abandoned. The 1:3 result suggests there is quality to recover if memory is less of a constraint, and the report promises checkpoints from the ablation runs so other people can look.

What is not yet shown

The model was trained on 8 GPUs with 80GB each, which is small by frontier standards, and the benchmark comparisons are against Mixtral and Llama-2 70B rather than anything larger. The 256K context claim rests on the report's own long-context evaluations. What has not been shown is whether the hybrid holds its ground at 10 or 20 times the training compute, or whether the attention share needs to grow as the model does.

The experiment we would want to see is the same architecture at a fixed budget with the attention ratio swept from 1:1 to 1:15, measured on both perplexity and a needle-in-a-haystack retrieval test at 256K. If retrieval quality degrades before perplexity does, that would tell us the attention layers are doing something specific at long range, and it would say where the floor on the ratio is.

Sources

  1. Lieber et al., Jamba: A Hybrid Transformer-Mamba Language Model (arXiv 2403.19887)