The part of the stack nobody upgraded

Yesterday Answer.AI and LightOn released ModernBERT, an encoder-only model in two sizes, 149 million and 395 million parameters, trained on 2 trillion tokens with a context length of 8,192. The release post opens with a number that explains why this matters: BERT, released in 2018, is still the second most downloaded model on the Hugging Face hub at more than 68 million downloads a month, and encoder-only models together account for over a billion downloads a month.

Those downloads are the retrieval and classification workloads that sit underneath most deployed language model systems. Every RAG pipeline that embeds documents, every reranker, every spam or abuse classifier that has to run on millions of items an hour is more likely to be an encoder than a decoder, because encoders are small and cheap and see the whole input at once. And almost all of them are running an architecture that predates rotary embeddings, gated activations, Flash Attention and the long-context training recipes that decoders have had for two years.

What changed in the architecture

Most of ModernBERT is a list of things decoder people will recognise. Rotary positional embeddings replace learned absolute positions, which is what allows the 8,192 context. GeGLU replaces GeLU in the feed-forward layers. Bias terms are removed. An extra normalisation layer is added after the embeddings for training stability. The next-sentence-prediction objective is dropped, and the masking rate is raised from 15 to 30 percent.

The two changes that are specific to the encoder setting are attention layout and padding. Attention alternates: every third layer is global over the full sequence and the rest use a local window of 128 tokens. That keeps the long-context cost from scaling quadratically at every layer. Unpadding and sequence packing remove the wasted compute on padding tokens in variable-length batches, which in a classification workload with short and long inputs mixed together can be most of the compute.

The authors also designed for the hardware people actually have. The model dimensions were chosen for consumer and inference GPUs, the RTX 3090 and 4090, the A10, the T4 and the L4, rather than for training clusters. That is a different design target from any recent decoder release and it is the right one for an encoder, since these models live in serving infrastructure rather than in research runs.

The training recipe

Training ran in three phases. The first 1.7 trillion tokens were at a sequence length of 1,024. Then 250 billion tokens at 8,192 to extend the context. Then 50 billion tokens of annealing on a differently sampled mixture. The data is described as mostly unique web documents, code and scientific articles, which is a deliberate departure from the Wikipedia-and-books mixture of the original BERT and the reason the model is competent on code. The large model was initialised by tiling the base model's weights rather than from scratch.

The numbers on the base model are good rather than spectacular on the classic benchmark and strong where it counts for this class of model. GLUE is 86.7 against 86.3 for DeBERTaV3-Base, and the post says this is achieved with less than a fifth of DeBERTa's memory. On BEIR retrieval it scores 54.7 against 53.4 for DeBERTaV3 and 53.8 for NomicBERT, though GTE-en-MLM is ahead at 56.8. On the long-document MLDR retrieval set it scores 48.9. On CodeSearchNet it scores 79.0 and on the StackQA code question set 80.2, where the post notes older encoders trained without code cannot score at all.

Speed is the headline

The number we would show a platform team is throughput. On an RTX 4090 with variable-length inputs, ModernBERT-Base processes 8.2 thousand tokens a second at a batch size of 128, against 2.1 thousand tokens a second at batch size 32 for DeBERTaV3-Base. The post claims roughly four times the speed of comparable encoders on mixed-length inputs and two to three times on 8,192 token inputs. If your reranker is the bottleneck in a retrieval pipeline, a four times throughput gain at equal or better quality is a direct cost cut.

The long context matters for a different reason. A 512 token encoder forces you to chunk documents before embedding them, and chunking is where a lot of retrieval quality quietly goes. An 8,192 token encoder lets you embed a whole page, a whole function, or a whole email thread as one unit and rerank on the full text. ColBERT-style late interaction models built on ModernBERT are reported at 9 points above other long-context models on the post's evaluation, which is the kind of gap you would expect when the model sees the document instead of a slice of it.

What we would check before swapping it in

Three things. First, these are base models with a masked language modelling objective, so the retrieval numbers come from fine-tuning them as embedders, and anyone with a tuned production embedder will need to redo that tuning on the new trunk. Second, the model ships in transformers 4.48 and drops the token_type_ids input that older BERT code passes, so existing pipelines will break in small ways. Third, the benchmark comparisons are the authors' own runs, and we would want the BEIR and MLDR numbers reproduced by someone with no stake before treating the retrieval gains as settled.

With those caveats, this is the release we have wanted for a while. The encoder side of the stack has been running on 2018 architecture because nobody with the compute cared enough to replace it, and the fact that it took a small lab and a sponsor's compute says something about where the field's attention has been. We expect the next year of retrieval papers to use this as the default trunk, and we expect the chunking heuristics half of us maintain to start disappearing.

Sources

  1. Hugging Face blog, Finally, a Replacement for BERT: Introducing ModernBERT