What Google shipped

On April 18 Google released quantization-aware trained versions of all four Gemma 3 sizes. The memory numbers are the point. Gemma 3 27B goes from 54 GB in bf16 to 14.1 GB in int4, 12B from 24 GB to 6.6 GB, 4B from 8 GB to 2.6 GB and 1B from 2 GB to 0.5 GB. Those figures are for weights only and the KV cache is extra, but 14.1 GB puts the 27B model on a single RTX 3090 with 24 GB, and 6.6 GB puts the 12B on an 8 GB laptop GPU such as the RTX 4060.

The training detail is short and specific. Google ran about 5,000 additional steps of training with the quantization simulated in the forward pass, using the probabilities from the unquantized checkpoint as the targets. That is distillation from the model into a quantized copy of itself. The claimed effect, measured in llama.cpp with Q4_0, is that the perplexity degradation from quantization fell by 54 percent compared with quantizing the ordinary checkpoint after the fact. The checkpoints are on Hugging Face and Kaggle and run in Ollama, LM Studio, MLX, gemma.cpp and llama.cpp.

Why QAT works where PTQ hurts

Post-training quantization takes a finished model and rounds its weights, then optionally patches up the damage with calibration data. The weights were never trained to sit on a coarse grid, so some fraction of them round badly and the model gets a little worse. At 8 bits nobody notices. At 4 bits the loss is measurable and on smaller models it is large.

Quantization-aware training changes what the optimiser sees. During those 5,000 steps the forward pass uses the rounded weights, so the gradient pushes the underlying full precision weights toward values that round well and that keep the output distribution close to the teacher. The model learns to be a good int4 model rather than being a good bf16 model that happens to be rounded. The cost is compute and access to a training pipeline, which is why until now most int4 checkpoints in the wild came from PTQ tools run by third parties. Google doing this in house for a whole family and publishing it is the mainstreaming that the title refers to.

The lossless route

The DFloat11 paper from Tianyi Zhang, Anshumali Shrivastava and colleagues at Rice and elsewhere, posted on April 15, starts from an observation about the bf16 format rather than about the model. A bf16 number has one sign bit, eight exponent bits and seven mantissa bits. The authors measure the entropy of the exponent across LLM weights and find it carries about 2.6 bits of information, because only about 40 of the 256 possible exponent values ever appear. The other six bits or so are structurally wasted.

DFloat11 Huffman codes the exponent and leaves the sign and mantissa untouched, which brings the average weight to a little under 11 bits. The results are strikingly uniform. Llama 3.1 8B goes from 16.06 GB to 10.90 GB, Qwen 3 14B from 29.54 GB to 20.14 GB, and Llama 3.1 405B from 811.71 GB to 551.22 GB, all around 68 percent of original size and around 10.8 to 10.9 bits per parameter. Since the decoding reverses the coding exactly, the outputs are bit for bit identical to the bf16 model. There is no accuracy question to argue about.

The engineering is in the decompression kernel. Weights are decoded on the GPU just before use, with the Huffman tree broken into lookup tables small enough for SRAM, a two phase kernel that first counts decoded elements and then writes them at computed offsets, and decompression batched at the level of a transformer block. The decode overhead is a fixed cost per forward pass, so it is amortised as batch size grows. The headline application is a 405B model, 810 GB in bf16, running on one node of eight 80 GB GPUs where it otherwise would not fit.

Choosing between them

These two results are often lumped together as compression, and they answer different questions. DFloat11 answers, can we keep the exact model and fit it in 30 percent less memory. QAT answers, can we accept a small controlled loss to cut memory by nearly four times. If your constraint is that outputs must match the reference model, for a reproducibility study, a certified deployment or a comparison where you need to rule out quantization as a confound, the lossless route is the only one that gives you a clean answer, and 30 percent is often the difference between fitting on a node and not.

If your constraint is consumer hardware, 30 percent is nowhere near enough. The 27B Gemma at 70 percent of 54 GB is still 38 GB, which fits nothing you can buy at retail. QAT at int4 gets it to 14.1 GB, and the 54 percent reduction in perplexity loss is the reason to prefer the official checkpoint over a community PTQ conversion. The right comparison for a practitioner is between Google's QAT int4 and the best PTQ int4 of the same model on the tasks you care about, and the blog post does not give that beyond the perplexity number. We would want to see it on a few downstream evals before trusting the 54 percent to generalise.

The two are also not exclusive. DFloat11 is described for bf16, but the same entropy argument applies to any format whose exponent or code distribution is skewed, and it is worth asking whether int4 QAT weights have enough structure left to compress further. Our guess is not much, because quantization already removes the redundancy that entropy coding exploits. That is a cheap experiment to run.

What we would test next

The DFloat11 throughput figures against CPU offloading, 2.31 to 46.24 times lower latency, are impressive but easy to win, since offloading is dreadful. The number we want is the latency against an uncompressed bf16 model on the same GPU at batch sizes of one and eight, which is what someone serving a model that already fits would pay to save memory. The paper says the overhead is constant and amortised, and the small batch case is where that argument is weakest.

For QAT, the experiment we would run is to apply the same 5,000 step distillation recipe to an open model of similar size where the base training code is available, and check whether the 54 percent figure reproduces. If it does, every open model release should ship with an int4 checkpoint trained this way, and the third party PTQ ecosystem becomes a fallback rather than the default.

Sources

  1. Google Developers Blog: Gemma 3 QAT Models: Bringing state-of-the-art AI to consumer GPUs
  2. arXiv: 70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (Zhang et al.)