The problem both methods solve

A 175 billion parameter model in 16-bit weights needs about 350 gigabytes just to sit in memory. Nobody has that on one GPU. Post-training quantization takes a trained model, rounds its weights to a smaller integer format, and hopes the output does not change much. The naive version, round-to-nearest, works acceptably at 8 bits and gets ugly below that. The question that defined the last eight months is how to get to 4 bits, and then 3, without retraining.

Two papers dominate the answer. GPTQ, from Frantar, Ashkboos, Hoefler and Alistarh at IST Austria and ETH, appeared in October 2022 and went to ICLR 2023. AWQ, from Song Han's group at MIT with Ji Lin as first author, was posted on June 1. They agree on the goal and disagree on almost every design decision, and the disagreement is instructive.

GPTQ: pay for the error with the other weights

GPTQ descends from Optimal Brain Surgeon by way of Optimal Brain Quantization. The idea is second-order. When you round one weight, you incur an error, and you can compensate by adjusting the weights that have not been quantized yet, using the inverse Hessian of the layer's reconstruction loss to decide how. OBQ did this one weight at a time in a greedy order, which does not scale past a few hundred million parameters.

GPTQ makes it scale with three changes. It quantizes columns in an arbitrary fixed order rather than searching for the best one, which the authors found costs little at scale. It batches the Hessian updates lazily so the GPU does useful work instead of memory traffic. And it reformulates the update through a Cholesky decomposition, which fixes the numerical instability that otherwise accumulates over billions of updates. The calibration set is 128 sequences from C4.

The results are why everyone uses it. OPT-175B goes from 12.51 perplexity on WikiText2 in FP16 to 15.45 at 4 bits and 24.18 at 3 bits. BLOOM-176B goes from 24.59 to 25.96 at 4 bits. Quantizing the 175B model takes around four GPU hours, and the paper reports this was the first time a 175B model ran generative inference on a single GPU. Measured speedups were 3.25 times over FP16 on an A100 and 4.5 times on an A6000. The paper is candid that small models are harder: the accuracy loss at a given bit width grows as the model shrinks.

AWQ: find the one percent that matters

AWQ starts from an observation rather than an optimizer. Some weights matter far more than others, and the way to find them is to look at the activations that flow through them rather than at the weights themselves. Table 1 of the paper makes the case on OPT-6.7B at 3 bits. Round-to-nearest gives 23.54 perplexity. Keeping one percent of channels in FP16 selected by weight magnitude gives 22.37, essentially nothing. Keeping one percent selected at random gives 24.23. Keeping one percent selected by activation magnitude gives 11.39.

Mixed precision is awkward for kernels, so the paper turns the insight into a rescaling. Multiply a salient input channel's weights by a factor s and divide the corresponding activation by s, and the layer computes the same function while the quantization error on those weights shrinks. The scale is set per channel as the average activation magnitude raised to a power alpha, and alpha is found by a grid search of 20 points on the interval from 0 to 1. There is no backpropagation and no reconstruction, just activation statistics collected offline.

On LLaMA-7B, FP16 perplexity is 5.68. AWQ at INT4 with group size 128 gives 5.78, and at INT3 with group size 128 gives 6.35 against 7.01 for round-to-nearest. The authors report a 3 times speedup over the Hugging Face FP16 implementation on desktop and mobile GPUs. The number we find most interesting is about calibration data. When the calibration set and the evaluation set come from different domains, PubMed against Enron for instance, AWQ's perplexity rises by 0.5 to 0.6 while GPTQ's rises by 2.3 to 4.9. AWQ also matches GPTQ with a calibration set ten times smaller. The second-order method fits its calibration data, and that shows up when the data moves.

What activation-awareness buys and what it does not

The trade is clear once the two are side by side. GPTQ spends compute to minimize a specific reconstruction loss on specific data, and gets the best number when the data matches. AWQ spends almost nothing and protects the channels the model actually uses, and holds up better when the data changes. For a model that will be deployed on inputs unlike C4, the second property matters more than a tenth of a point of perplexity.

They are also not exclusive. The AWQ paper runs the combination at INT2 with group size 64 on OPT-6.7B. Round-to-nearest collapses to a perplexity of 7,622. GPTQ alone reaches 16.65. AWQ scaling followed by GPTQ reaches 15.71. So the scaling transformation is orthogonal to the second-order update and can be applied first. Below 3 bits, though, look at those numbers again. From 11 or 12 at 3 bits to 16 at 2 bits is a large jump for the model to absorb, and it is the cliff every team we know has hit trying to go lower.

Rules we would write down

First, 4 bits with grouping is free for models above about 7B and nearly free at 3 bits with either method. Second, below 3 bits, expect a step change rather than a slope, and do not trust a single perplexity number to tell you where the step is. Third, if you use GPTQ, calibrate on data that looks like production, because the method will overfit whatever you give it. Fourth, measure downstream accuracy and not just WikiText2, since neither paper claims perplexity is the metric users care about.

The open question we would like someone to run is whether the one percent of salient channels is stable across tasks. If the channels that matter for code are the channels that matter for prose, AWQ's scaling is a property of the model. If they differ, it is a property of the calibration set, and the cross-domain result deserves a closer look than a single cross-domain experiment.

Sources

  1. Frantar et al., GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers (arXiv 2210.17323)
  2. Lin et al., AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration (arXiv 2306.00978)