Mamba was rejected from ICLR, and everyone had an opinion about peer review
The most discussed architecture paper of the last year came back from ICLR 2024 with scores of 8, 8, 6 and 3 and a rejection. The reviewers wanted Long Range Arena results and doubted perplexity as the headline metric. Both of those are reasonable, and the argument that followed was about something else.
What happened
Albert Gu and Tri Dao posted Mamba on December 1, 2023. It replaces attention with a selective state space model whose parameters depend on the input, so the model can propagate or forget information along the sequence depending on the current token, and it comes with a hardware-aware parallel algorithm to make that practical. The headline claims are 5 times the throughput of a transformer, linear scaling in sequence length, and a 3B model that outperforms transformers of the same size and matches transformers twice its size.
Within weeks the paper was everywhere. Implementations, ports, follow-up architectures, reading groups. Then in late January the ICLR 2024 decisions came out and it was rejected. Sasha Rush posted the reaction that got quoted the most: Mamba apparently was rejected, honestly we do not even understand, if this gets rejected what chance do the rest of us have.
The scores, as reported in the write-ups circulating since, were 8, 8, 6 and 3. Three reviewers were positive to enthusiastic and one was not, and the decision followed the low score. We could not open the review thread directly while writing this, so we are relying on secondhand accounts of the reviews and would encourage anyone quoting the specifics to check the thread themselves.
The objections were not silly
Two criticisms come up in every summary. The first is the absence of Long Range Arena results. LRA is the standard benchmark for long-sequence models, and the state space model line the paper descends from was built and validated on it. Asking a paper that claims linear-time long-sequence modeling to report the long-sequence benchmark is exactly the kind of request reviewers exist to make.
The second is scepticism about perplexity as the primary evidence. Lower perplexity does not reliably mean better performance on the tasks people care about, and a new architecture claiming parity with transformers on language modeling loss has not yet shown parity on anything a user would notice. The paper does report downstream zero-shot evaluations, and a reviewer can still reasonably say that a model matching transformers on language modeling loss has not yet been stress tested on the tasks that would expose a missing mechanism.
So the reviewers asked for the right things. What they got wrong, if anything, was the weight. A paper with a working implementation, a large reported speedup and a clear mechanism is a contribution even with a benchmark missing, and the fix for a missing benchmark is a revision request.
What the argument was actually about
The reason this became a public fight has little to do with Mamba. By the time the decision landed, everybody who cared had read the paper, run the code, and formed a view. The reviewers were being asked to certify something the community had already evaluated in public, at a scale and speed the review process cannot match. Their verdict arrived as an opinion about a fact that was already settled elsewhere.
That is the structural change worth naming. When a paper appears on arXiv with code and gets reproduced within a fortnight, peer review stops being the mechanism that decides whether the work is real. It still does other things. It checks whether the claims match the experiments, it demands the comparison the authors skipped, and it puts a stranger with domain knowledge in the loop before a result enters the literature. Those are valuable and they are not the same as gatekeeping.
The people arguing that the rejection proved review is broken and the people arguing that the review was correct were mostly not disagreeing. A single reviewer at 3 can sink a paper with two 8s, which is a fact about how decisions aggregate scores rather than about the quality of any review. That is a conference process question and it has known fixes, including requiring an area chair to justify a decision against the majority of reviewers.
What we take from it
Publish the reviews. ICLR already does, which is the only reason this conversation could happen with any specifics at all, and it is why we can point at the LRA complaint instead of guessing. Every venue that keeps its reviews private converts an argument about evidence into an argument about vibes.
And then treat acceptance as what it is, one signal among several, weaker than reproduction and much weaker than a year of people building on the work. Mamba will be cited thousands of times whether or not it appeared in the ICLR proceedings, and the field will get its long-sequence numbers from someone eventually. The cost of the rejection falls mostly on the review process itself, which spent its credibility on a verdict nobody was waiting for.
Sources
From the foundation