Byte Latent Transformer: patches instead of tokens
Meta's BLT drops the tokeniser and groups raw bytes into patches whose boundaries are set by next-byte entropy. At 8B parameters and 1T tokens of data it matches or beats a Llama 3 baseline on average while offering a way to trade accuracy for up to 50 percent fewer inference flops. Notes on the design and what the tables actually show.
What is being replaced
Nearly every language model in production is trained on subword tokens produced by a BPE tokeniser that was built before training started and never updated. The tokeniser decides how much text each transformer step covers, and it decides that uniformly. A common word costs one step and so does a rare, hard-to-predict one. Meta's Byte Latent Transformer paper, out this week from Artidoro Pagnoni and thirteen colleagues, is the most serious attempt we have seen to remove that component at scale.
The replacement unit is a patch, a variable-length run of raw bytes. Patch boundaries are chosen by a small byte-level language model trained on the same data distribution. When that model's entropy for the next byte crosses a global threshold, a new patch begins. The paper's own illustration runs the sentence about Daenerys Targaryen through the patcher and the point lands cleanly. The first byte of a name is uncertain and starts a patch, and the rest of the name is low entropy so the patch just extends across it. Compute goes where prediction is hard.
How the architecture is arranged
BLT has three pieces. A lightweight local encoder turns bytes into patch representations, using hash n-gram embeddings over byte n-grams of size 3 to 8 plus cross attention. A large latent global transformer runs once per patch and does most of the work. A local decoder then unpacks each patch prediction back into bytes. The global model is the expensive part, so the number of patches per document, which is the average patch size, is the main lever on flop cost.
This is the trick that makes the scaling argument possible. With a fixed vocabulary tokeniser you cannot simply ask for bigger tokens. With entropy patching you set a threshold and get whatever average patch size you want, and the paper states that Llama 3's own tokeniser only raised average token size from 3.7 to 4.4 bytes by growing the embedding table four times over compared with Llama 2. BLT-Entropy at patch size 8 covers nearly twice as many bytes per step as that.
The scaling study
The headline is a flop-controlled scaling study up to 8B parameters and 4T training bytes, the first at this size for byte-level models. The interesting plots hold inference flops fixed rather than training flops. Table 2 lists the matched model sizes: at the small class, Llama 3 with BPE gets 450M non-embedding parameters while BLT-Entropy at patch sizes 6 and 8 gets 610M and 760M for the same inference cost per byte. At the larger class the numbers are 3.9B against 5.2B and 6.6B. Longer patches mean the global model runs less often, so the saved compute buys a bigger model.
The BPE models win at small training budgets and BLT overtakes them shortly past the compute-optimal point. The paper reports the crossover at roughly 150B bytes for the small class and 1T bytes for the larger one, about 2.5 to 3 times the compute-optimal data quantity. That matters because the models people actually deploy, Llama 3.1 8B being the example the authors pick, are trained on two orders of magnitude more data than compute-optimal. In that regime the trend lines favour patches.
The 8B head-to-head against Llama 3
Table 1 is the one we would send a sceptical colleague. Three flop-matched 8B models trained on the same 1T token dataset: a Llama 3 tokeniser baseline, BLT-Space with boundaries at spaces, and BLT-Entropy. On the seven-task average BLT-Entropy scores 61.1 against Llama 3's 60.0 and BLT-Space's 58.0. It wins on Arc-Easy (79.6 vs 77.6), HellaSwag (80.6 vs 79.1), MBPP (41.8 vs 40.2) and HumanEval (35.4 vs 31.1), and loses a little on Arc-Challenge, PIQA and MMLU (57.4 vs 58.1).
Two caveats sit in the table's own footnotes. BLT-Entropy's average patch size on the training mix is 4.5 bytes, almost identical to the 4.4 bytes of the Llama 3 tokeniser, so this comparison shows parity at equal granularity rather than a flop saving. The saving comes from BLT-Space at 6.1 bytes per patch, which costs about two points of average accuracy in exchange. The paper also notes that for the entropy model they lowered the threshold at inference time from 0.6 to 0.1, which improved task scores at the cost of more steps. The 50 percent figure in the abstract is a trade you can choose to make, not something you get for free.
Where bytes clearly win
Robustness is where the gap opens. On noised HellaSwag, where the text is dropped, uppercased, repeated or written as space-separated characters, the 1T token Llama 3 averages 56.9 and BLT averages 64.3, matching Llama 3.1 which saw 16 times more data. On the CUTE character-understanding suite BLT scores 54.1 against 27.5 for Llama 3 and 20.0 for Llama 3.1. The spelling subtask is the extreme case, 99.9 for BLT and 1.1 for the tokenised model. Anyone who has watched a frontier model fail to count the letters in a word will recognise that number.
The authors' reading, which we share, is that byte-level awareness is not something more data buys you. Llama 3.1 with 16T tokens does worse than Llama 3 with 1T on several CUTE subtasks. The tokeniser is hiding characters from the model, and no amount of training recovers what the input pipeline threw away.
Why this is still hand-engineered, and what we want to see
It is worth being precise about what BLT removes. The BPE vocabulary is gone, but the entropy model and its threshold are new hand-chosen components, and the paper spends real effort on details like resetting the entropy model's context at newlines to stop patches from drifting larger inside repetitive structured text. The segmentation decision has moved from a frequency table to a small learned model, which is progress, but the boundary is still fixed before the main model trains.
What we would want next is the obvious experiment the paper cannot yet run: train the entropy model and the latent transformer together so patching adapts to the model that consumes it. We would also like the fixed-inference scaling curves pushed past 8B, since the whole case for patches rests on the crossover moving in BLT's favour as models grow. If the trend holds one more order of magnitude, the tokeniser stops being infrastructure and becomes a legacy decision.
Sources
From the foundation