1.58 bits: what BitNet promised and what it required
Microsoft's BitNet b1.58 trains a language model whose every weight is minus one, zero or one, and claims parity with FP16 at 3B parameters. The catch is in the word trains.
Three values per weight
A paper from Shuming Ma, Hongyu Wang, Furu Wei and colleagues at Microsoft went up on arXiv at the end of February with a title that reads like a manifesto. The Era of 1-bit LLMs. The model it describes, BitNet b1.58, restricts every weight in its linear layers to one of three values, minus one, zero or plus one. Three states need log2(3) bits of information, which is 1.58, hence the name.
The quantisation function is simple. Take the weight matrix, divide by its mean absolute value, round each entry to the nearest of minus one, zero and one. They call it absmean. Activations are kept at 8 bits and scaled per token, without a zero point. The zero is the important addition over their earlier binary BitNet. It lets the model switch a weight off entirely, which the authors frame as a form of feature filtering, and it is what pushes ternary past binary in quality.
The reason the claim caused such a stir is that with ternary weights, matrix multiplication no longer needs multiplication. Every product of a weight and an activation is either the activation, its negation or nothing, so the whole operation collapses to integer addition. The paper reports that this cuts arithmetic energy for the matrix multiply by 71.4 times on a 7nm process compared with FP16. That is the number people passed around.
The parity claim, with its numbers
The headline is that a ternary model matches a full-precision one of the same size trained on the same data. The evidence is a set of LLaMA-architecture models from 700M to 3.9B parameters trained on 100B tokens of RedPajama, compared against FP16 LLaMA trained identically.
At 700M, BitNet's perplexity is 12.87 against 12.33 for FP16, and its average across seven zero-shot tasks is 44.3 percent against 45.5. At 1.3B the gap narrows to 11.29 versus 11.25 perplexity and 45.4 versus 46.2 percent. At 3B the ternary model pulls ahead, 9.91 versus 10.04 perplexity and 50.2 versus 49.7 percent. A 3.9B ternary model reaches 9.62 perplexity and 51.2 percent, which is better than the 3B FP16 baseline while using less memory.
Read that trend carefully. The gap closes with scale, which is the pattern you want if you believe large models have more redundancy to absorb the quantisation. But the largest full comparison is 3B parameters on 100B tokens, which is a small run by current standards. There is a separate result at 2T tokens, where a 3B BitNet averages 74.34 percent across five tasks against 73.22 for StableLM-3B, but that is a comparison against a different model trained by a different group, not a controlled pair.
What you get on today's hardware
The efficiency numbers were measured on ordinary GPUs, which is worth stating because the paper's rhetoric is about future chips. At 3B, BitNet b1.58 uses 3.55 times less GPU memory than FP16 LLaMA and runs 2.71 times faster in latency. The authors extrapolate to 70B and report 4.1 times lower latency, 11 times larger maximum batch size and 8.9 times higher throughput, on the assumption that memory bandwidth is the limit.
Those gains come almost entirely from moving fewer bytes. On current hardware the weights are still unpacked and multiplied like any other low-bit model, so the 71.4x energy figure is a paper calculation for a matrix unit that does not exist yet. The authors say so, and the last section of the paper argues explicitly for hardware designed around 1-bit models, comparing the opportunity to what Groq has done with its LPUs. That is the promise. The requirement is that someone builds the chip.
Why post-training quantisation is still the practical route
The word that matters in the paper is trained. BitNet b1.58 is quantisation-aware from the first step. The model never has full-precision weights to fall back on. Every forward pass uses the ternary weights, and the model learns to be good under that constraint. That is a very different proposition from taking a checkpoint you already have and compressing it.
Most of us do not train models from scratch. We download them. For a downloaded Llama or Mistral checkpoint the choice is post-training quantisation, methods like GPTQ and AWQ that squeeze a trained model to 4 bits with a small calibration set and a few GPU hours. They do not reach 1.58 bits and they lose a little quality, but they work on any model that exists today, and they work this afternoon. BitNet's recipe requires a lab with the compute to pretrain, and a willingness to bet that compute on an architecture whose largest published comparison is 3B parameters.
There is also a data question the paper does not answer. If ternary models match FP16 at equal tokens, do they keep matching at 10 trillion tokens, where current open models are heading? Redundancy is what lets a model absorb quantisation, and long training reduces redundancy. We do not know which way that goes, and we would want to see a 7B ternary model trained to a modern token budget before we believed the era had begun.
What we would want next
The result we would most like to see is not another benchmark table. It is a released checkpoint. The paper reports numbers and describes the method, but there is no ternary model to download and no kernel to run it with. The whole argument for 1.58 bits is that it changes what inference costs, and that argument is only testable if someone outside Microsoft can measure it.
Until then, the useful reading of BitNet b1.58 is as an existence proof. A model with three-valued weights can be trained to the same quality as a full-precision one at small scale. Whether that becomes the way models are built depends on hardware that has not been designed, training runs that have not been paid for, and a scaling question that has not been asked above 4B parameters. Post-training quantisation will stay the default for a while, and the ternary result is the strongest reason yet to think the default might eventually change.
Sources
From the foundation