LoRA learns less and forgets less
Reading notes on the Databricks comparison of LoRA and full finetuning for Llama-2-7B on code and math. Adapters lag badly in continued pretraining, nearly catch up in instruction tuning at high rank, and keep more of the base model either way.
The setup
Dan Biderman and colleagues at Databricks posted a paper on May 15 that does the comparison most of us have been doing informally and never publishing. They take Llama-2-7B and finetune it on two target domains, programming and mathematics, in two data regimes. Instruction finetuning uses roughly 100K prompt-response pairs (Magicoder-Evol-Instruct-110K for code, MetaMathQA for math). Continued pretraining uses up to 20B unstructured tokens (StarCoder-Python and OpenWebMath). They run LoRA at ranks 16, 64 and 256 against full finetuning, with a learning rate sweep for each method.
Learning is measured on HumanEval (164 problems, pass@1) and GSM8K (the 1,319 problem test split). Forgetting is measured as the average of HellaSwag, ARC-Challenge and WinoGrande, the standard Open LLM Leaderboard tasks, so it captures how much general capability the model loses while gaining the target skill. We like that both axes are reported on the same runs. Most LoRA papers report only the first.
Where LoRA underperforms
In continued pretraining the gap is large and it does not close. On code, the best LoRA model at rank 256 peaks at a HumanEval of 0.224 after 20B tokens. Full finetuning reaches 0.218 at 4B tokens and 0.263 at 20B. So the adapter needs five times the data to match what full finetuning does early, and never reaches where full finetuning ends up. Math continued pretraining looks the same. LoRA at rank 256 peaks at GSM8K 0.203 at 16B tokens, below full finetuning at 4B tokens (0.224) and far below its 20B peak of 0.293.
Instruction finetuning is friendlier to adapters. On math, rank 64 gets to GSM8K 0.624 at epoch 4 and rank 256 peaks at 0.634, against full finetuning's 0.642. On code the ordering by rank is visible from the first epoch and only rank 256 gets close to full finetuning. The authors' explanation is that the math data is mostly English with a bit of arithmetic, so it sits near the pretraining distribution, while code is far enough from it that a low rank perturbation cannot carry the change.
Where LoRA wins
The forgetting numbers flip the picture. In code continued pretraining at 20B tokens, full finetuning drops the forgetting average to 0.545 while LoRA at rank 256 holds 0.617. In code instruction tuning, full finetuning is at 0.414 versus 0.509 for LoRA at rank 64. On math the two methods forget about equally, which matches the learning story, since there was less to move in the first place.
The more interesting comparison is against ordinary regularisation. On the Magicoder data they compare LoRA to weight decay at 5e-5 and 1e-4 and to attention dropout at 0.05 and 0.1. The regularised full finetuning runs learn and forget roughly like unregularised full finetuning. LoRA at rank 16 learns less and forgets less than everything. LoRA at rank 256 learns as much as the full finetuning variants and still forgets less. If your goal is keeping the base model intact, the adapter does something dropout and weight decay do not.
They also count unique outputs over 50 generations per HumanEval problem and find that full finetuning collapses generation diversity more than LoRA does. We would treat that as a side observation rather than a headline, but it points the same direction.
Why the gap exists
The section we keep returning to is the spectral analysis. They take the difference between finetuned and base weight matrices and ask what rank is needed to explain 90 percent of its variance. For full finetuning on code the answer is 10 to 100 times the ranks people normally use for LoRA, and it grows as training continues. MLP modules end up higher rank than attention modules, and the first and last layers are lower rank than the middle. The premise that finetuning is a low rank update to the base weights, which is the argument for LoRA in the first place, does not hold on these datasets.
That also explains a practical finding. Targeting only attention matrices, which is how LoRA was introduced, underperforms targeting MLP modules or all modules, and most of the gain in the all-modules setting comes from the MLP blocks.
The recommendations, and the caveats
Their best practices are short. Use LoRA for instruction finetuning rather than continued pretraining. If memory allows, target all modules at rank 256, because ranks 16 to 64 tend not to suffice for code. Set the scaling factor alpha to twice the rank, since common libraries scale the update by alpha over r and quietly shrink high ranks. Sweep learning rates between 1e-5 and 5e-4 and take the highest that trains stably, which usually lands an order of magnitude above the right value for full finetuning. They also note that at a fixed batch size LoRA trains slower than full finetuning in standard implementations, which surprised us.
The caveats are the obvious ones. One base model at 7B, two domains, and benchmarks that reward exact answers. We do not know whether the rank gap shrinks on a larger model where the base already knows most of the target domain, and the paper does not claim to. What we take from it is narrower and still useful. If you are using LoRA because it is cheap, check your target benchmark against a full finetuning run at least once, and if you are using it because you want the base model preserved, you are getting that, and getting it more reliably than from weight decay.
Sources
From the foundation