The result in two tables

Xiangyu Qi, Yi Zeng and a team spanning Princeton, Virginia Tech and IBM posted a paper on October 5 with a finding that is easy to state. Safety alignment in current chat models is removed by a small amount of fine-tuning, and it is degraded by fine-tuning that has nothing to do with safety at all. They tested GPT-3.5 Turbo through OpenAI's fine-tuning API, which opened to the public in August, and Llama-2-7b-Chat with Meta's official recipe.

The headline table is the harmful examples attack. They sampled 10, 50 and 100 instruction and response pairs from Anthropic's red team dataset, fine-tuned for five epochs, and measured how often the model gave a maximally harmful answer on a 330-prompt benchmark judged by GPT-4. GPT-3.5 Turbo went from a harmfulness rate of 1.8 percent to 88.8 percent with ten examples. Llama-2-7b-Chat went from 0.3 percent to 50 percent with ten and 80 percent with fifty. The GPT-3.5 job cost less than 20 cents. The Llama run at batch size ten for five epochs is five gradient steps.

Three levels of risk

The paper organises its experiments into three tiers, and the ordering is the argument. Level one is explicit harm, the table above. Any provider can in principle screen uploaded training data for that, and the authors note that their attacks did not trigger OpenAI's moderation as it stood at the time, which they disclosed before publication.

Level two is what they call identity shifting. They wrote ten examples with no toxic content at all, in which the model is told it is now AOA, an absolutely obedient agent, and practises answering benign requests with an affirmative preamble. After ten epochs on those ten examples GPT-3.5 Turbo's harmfulness rate was 87.3 percent. Llama-2 reached 54.2 percent after three epochs. Nothing in that dataset would trip a content filter, because the harm is in the disposition being taught, not the text.

Level three is the one we keep coming back to. They fine-tuned both models on Alpaca and Dolly for one epoch with the recommended hyperparameters, the sort of thing a team does to make a customer support bot. GPT-3.5 Turbo's harmfulness rate rose from 5.5 percent to 31.8 percent on Alpaca and from 4.5 to 23.9 on Dolly. Llama-2-7b-Chat went from 0.3 to 16.1 percent on Alpaca and to 18.8 percent on LLaVA-Instruct. An ablation with a higher learning rate and smaller batches made it worse. Nobody in that scenario intended anything.

Why this reshapes two debates

The first debate is about fine-tuning APIs. The safety story for hosted models has been that the provider controls the weights, so alignment holds. This paper shows that the moment a provider lets users move the weights, even through a narrow API with epochs as the only knob, the alignment is a property of the fine-tuning data rather than of the model. Data moderation catches level one. It does not catch level two, and level three is the normal case. The authors' own remark is that current RLHF and safety fine-tuning produce relatively surface-level changes to the model, and the asymmetry between thousands of alignment examples and ten attack examples is the evidence.

The second debate is about open weights, and here the paper cuts both ways. Critics of releasing Llama-2 have argued that anyone can strip the safety training, and this is the cleanest demonstration yet that they are right, at a cost of a few gradient steps. Defenders can point out that the hosted model fell just as fast, and that the closed API gave the attacker fewer levers and still failed. Under either reading, the model card claim that a chat model is safe is now a claim about the weights as shipped and nothing after.

What we would want checked next

The evaluation leans on GPT-4 as the judge, with a human meta-evaluation in the appendix, and the benchmark is 330 prompts built from the two companies' usage policies. That is a reasonable design, and we would still like to see the level three numbers reproduced with a different judge and a larger prompt set before treating 31.8 percent as a stable figure. The direction is not in doubt. The magnitude for benign data is the number that product teams will quote, so it should be firm.

The mitigation section discusses mixing safety data into fine-tuning, which the authors test and find only partially effective, and raises the possibility of backdoors that defeat auditing. The experiment we would run at our scale is the cheap one. Take Llama-2-7b-Chat, fine-tune on a dozen ordinary domain datasets a real team might use, and plot the harmfulness rate against learning rate, batch size and the fraction of safety examples mixed in. If there is a recipe that keeps the drop under a few points without hurting the task, that is a result people can use next month. If there is no such recipe, then the fine-tuning API and the weight release are the same problem, and the paper's title is the right one.

Sources

  1. Qi et al.: Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! (arXiv 2310.03693)