What TII actually trained

Falcon Mamba 7B is a causal decoder with 64 layers, hidden dimension 4,096, SSM state dimension 16 and a vocabulary of 65,024. There is no attention anywhere in it. It was trained on roughly 5,500 gigatokens, mostly RefinedWeb with technical, code and math data plus some FineWeb-Edu, on 256 H100 80GB GPUs across 32 AWS SageMaker p5 instances, in bfloat16, over about two months. The learning rate schedule was warmup-stable-decay with a maximum of 6.4e-4, batch size ramped from 128 to 2,048, and the context length was extended from 2,048 to 8,192 during training. The team added RMS normalization layers to the original Mamba block for stability at scale.

Two years of state space model papers had produced strong small models and one 7B, TRI-ML's mamba-7b-rw, that scored 45.52 on the legacy Open LLM Leaderboard. That number is the reason the field treated pure SSMs as a research direction rather than a deployment option. Falcon Mamba scores 64.09 on the same board: 62.03 ARC, 80.82 HellaSwag, 62.11 MMLU, 73.64 Winogrande, 53.42 TruthfulQA, 52.54 GSM8K.

The comparison that matters

On the legacy board, Falcon Mamba at 64.09 sits above Meta-Llama-3-8B at 62.62, Meta-Llama-3.1-8B at 62.28, Mistral-7B-v0.1 at 60.97 and Gemma-7B at 63.75, and just below Falcon2-11B at 64.28. Against the other recurrent and hybrid models it is not close: RecurrentGemma-9B scores 57.95, Zamba-7B-v1 scores 60.00, and mamba-7b-rw scores 45.52.

The newer v2 board, which is harder and weights instruction following and reasoning differently, tells a flatter story. Falcon Mamba averages 15.04 against Gemma-7B at 15.28, Mistral-Nemo-Base 12B at 15.08, Mistral-7B at 14.50 and Llama-3.1-8B at 13.78. Its IFEval score of 33.36 is the highest in that table by a wide margin, and its MMLU-PRO of 14.47 is among the lowest, with Llama-3.1-8B at 24.95. So the model follows instructions well and does worse on the knowledge-heavy multiple choice, which is roughly what we would expect from a recurrent model trained on a web-heavy mixture.

On efficiency, the blog measures generation with a prompt of length 1 and up to 130,000 generated tokens at batch size 1 on an H100, and reports constant throughput with no increase in CUDA peak memory. A transformer on the same experiment degrades as the KV cache grows. That is the structural claim of the architecture, demonstrated rather than argued.

What this settles

It settles the existence question. A pure state space model at 7B parameters, given a competitive data budget and a serious training setup, lands in the same band as contemporaneous transformers of similar size. Every previous attention-free release could be dismissed as undertrained or too small to compare. This one cannot, and that removes the easiest objection to the whole line of work.

It also settles a practical point about deployment. Constant memory during generation is not a benchmark artifact, and for long-generation workloads on constrained hardware it changes what fits. The model card notes it can process larger sequences than a comparable transformer on a single A10 with 24GB.

What it does not settle

In-context retrieval is the open weakness, and the evidence for that came from the Jamba paper in March rather than from anything in this release. The AI21 team ablated pure Mamba against pure attention at equal scale and found specific failures. On IMDB, pure Mamba scored 48.8 against 84.1 for attention. On QuAC, 20.2 against 27.9. On NarrativeQA, 27.7 against 45.8. The IMDB case is the diagnostic one. The task requires the answer Positive or Negative, and pure Mamba produced things like Very Good, Funny and 3/10. The model understood the sentiment and could not pick up the output format from the examples in its context.

Their explanation is that the lack of an attention mechanism makes in-context learning hard, and they connect it to induction heads, the attention circuits that copy patterns from earlier in the sequence. Mamba can be trained to copy, and what it does not do is develop the emergent in-context behaviour that shows up in transformers without anyone asking for it. Falcon Mamba's IFEval score suggests instruction tuning can cover a lot of this ground. It does not tell us what happens when the required behaviour appears only in the prompt.

That is why the hybrids won the architectural argument even as this release won the capability one. Jamba interleaves one attention layer per eight, keeps 256K context on a single 80GB GPU with 52B total and 12B active parameters, and reports 3 times the long-context throughput of Mixtral 8x7B. One attention layer in eight is cheap, and it buys back the thing the ablation shows pure recurrence lacks.

What we would test next

The experiment this release makes possible is a clean one. Run Falcon Mamba against a same-size transformer on tasks where the answer format, the label set or a required mapping exists only in the prompt, with no fine-tuning on either side. Vary the number of in-context examples. If the gap closes with more examples, the deficit is sample efficiency and it will shrink with scale. If the gap holds at 32 examples the way it held at a few, the induction head account is right, and the constant-memory advantage will keep costing something that only attention supplies.

We would also want the 130,000-token throughput result paired with an accuracy result at the same lengths. Constant memory is only useful if the model still finds what it needs at position 100,000, and neither the blog nor the model card reports a long-context retrieval evaluation. That is the number we would publish next if this were our model.

Sources

  1. Hugging Face: Welcome Falcon Mamba, the first strong attention-free 7B model
  2. Model card: tiiuae/falcon-mamba-7b
  3. Gu and Dao, Mamba: Linear-Time Sequence Modeling with Selective State Spaces (arXiv 2312.00752)
  4. Lieber et al., Jamba: A Hybrid Transformer-Mamba Language Model (arXiv 2403.19887)