Precision as a third axis

The standard scaling law has two knobs, parameters and tokens, and treats the number format as a fixed detail of the hardware. A paper posted on November 7 by Tanishq Kumar and eight coauthors, including Christopher Ré and Aditi Raghunathan, adds precision as a third knob and fits a law to it. The fits come from more than 465 pretraining runs of OLMo-style transformers on Dolma, at non-embedding sizes from 30 million to 220 million parameters and 1.5 to 26 billion tokens, and are validated on models up to 1.7 billion parameters.

The core idea is simple to state. Training in lower precision reduces the model's effective parameter count. A weight stored in fewer bits carries less information, so the model behaves like a smaller one, and the loss curve shifts accordingly. Everything else in the paper follows from putting a functional form on that intuition and checking it against data.

The fitted form

The loss is written as A times N_eff to the minus alpha, plus B times D to the minus beta, plus a constant, plus a post-training quantisation penalty. The effective parameter count is the true count multiplied by a factor of the form one minus e to the minus P over gamma, once each for the precision of weights, activations and the KV cache, with a separate fitted sensitivity gamma for each. At high precision the factor approaches one and you recover the ordinary law. As bits fall, the factor collapses and the model acts smaller.

The post-training quantisation term is where the practical result lives. The degradation from quantising a trained model scales as a constant times D to a power gamma_D, divided by N to a power gamma_N, times e to the minus P_post over gamma_post. Read that as a function of the tokens-per-parameter ratio. The more data a model of a given size has seen, the larger the penalty when you quantise it afterward, and the penalty grows without bound as D over N grows.

More data can make the quantised model worse

That last point is the headline and it is worth stating carefully. The authors observe that post-training quantisation degradation increases with training data size across all model sizes they tested. For a model that will be served in low precision, there is a D over N ratio beyond which additional pretraining tokens lower the full-precision loss but raise the quantised loss by more, so the deployed model gets worse. The experiments cover ratios up to roughly a thousand tokens per parameter, which is well inside the range where current small models are trained.

The mechanism the paper offers is that heavily overtrained models pack more into each weight and are therefore more sensitive to the noise quantisation adds. Whatever the mechanism, the operational consequence is that the choice to serve in low precision should be made before pretraining budget is set, since it changes the optimal amount of data.

Where the floor is

When parameters, data and precision are optimised together, the paper finds compute-optimal pretraining precision lands around 7 to 8 bits, and that this optimum is roughly independent of the compute budget. That is a striking claim because it says the industry's convergence on 8-bit formats is close to the right answer rather than a temporary stop on the way to 4 bits. When model size is held fixed, which is how model families are actually built, the optimal precision instead grows with the logarithm of compute, so bigger training runs at fixed size want more bits.

Below 4 bits the fits say the effective parameter count falls so fast that model size must grow by more than four times to hold loss scaling, which wipes out the compute saving. The paper also notes that its integer quantisation laws fit well down to 4 bits, at which point floating-point formats become noticeably more expressive and the two families diverge. So the floor is not a single number. It is a region around 4 bits where the accounting stops favouring fewer bits, and the exact edge depends on the format.

What the fits cannot tell you

The authors are careful about scope and their caveats matter for anyone tempted to plug these formulas into a training plan. The architecture is fixed throughout, whereas production low-precision training uses architectural changes for stability. Quantisation during training is applied only to the forward pass. The analysis is loss only, with no downstream evaluations. And the largest model in the study is 1.7 billion parameters, which is why they describe the trends as suggestive rather than prescriptive. There is also a systems point: compute cost scales linearly with precision in theory, and the real gain from halving precision is less than two times because of overhead.

What we take from the paper is a way to make the low-bit debate quantitative. The claim that FP8 is fine and FP4 is coming has been argued from anecdotes about particular runs. Here is a functional form that predicts when each will pay off, fitted on hundreds of runs, that someone with a larger budget can now try to falsify. The most useful follow-up would be to repeat the D over N experiment at 7 billion parameters with a modern format, and to report whether the point where more data starts hurting the quantised model moves, and in which direction.

Sources

  1. Kumar et al., Scaling Laws for Precision (arXiv 2411.04330)
  2. Full text (arXiv HTML)