Two ways safety training fails

Jailbreaks have so far been collected like folk remedies. Someone finds a prompt that works, posts it, and it circulates until the vendor patches it. The paper Alexander Wei, Nika Haghtalab and Jacob Steinhardt posted on July 5 is the first attempt we have seen to say why they work, and it reduces the zoo to two mechanisms.

The first is competing objectives. A safety-trained model is still trained to predict text and to follow instructions, and those objectives can be set against the safety objective. The second is mismatched generalisation. Pretraining covers a far wider range of inputs than safety training does, so there are inputs the model understands perfectly well that safety training never taught it to refuse. The two are different in kind. The first is a conflict inside the training signal. The second is a gap in its coverage.

Competing objectives, with examples

The cleanest example is prefix injection. The prompt asks the model to begin its answer with an innocuous-sounding prefix and then continue. Once the model has produced the prefix, continuing with a refusal would be an unnatural continuation of its own text, so the pretraining objective pulls it toward compliance. The authors ran an ablation. Change the required prefix to a plain Hello and the attack on GPT-4 stops working. The specific prefix is doing the work.

Refusal suppression is the instruction-following version. The prompt lists rules such as never apologise, never include a disclaimer, never use the word cannot. The model has been trained hard to follow formatting instructions, and a refusal violates all of them at once. The authors also place style injection and many of the popular role-play prompts, including the top prompt on jailbreakchat.com, in this category. They are all ways of making the safe answer cost more under one of the model's other objectives.

Mismatched generalisation, with examples

The signature example is Base64. Encode the harmful request, ask the model to respond in Base64, and GPT-4 will produce synthesis instructions it would refuse in plain English. The explanation the authors offer is that large models learn Base64 from pretraining and can follow instructions in it, while safety training data contains nothing that looks like Base64. The capability generalised to the encoding and the refusal did not.

The evidence that this is about capability comes from the smaller model. GPT-3.5 Turbo, given the same Base64 prompts, mostly says it cannot understand them. It has no vulnerability there because it has no capability there. GPT-4 has both. That asymmetry is the core observation of the paper and the reason scaling does not obviously help. A bigger model learns more encodings, more languages and more obscure formats, and each one is a new region the safety training has to reach.

The numbers

The evaluation uses a curated set of 32 harmful prompts drawn from the red-teaming reported by OpenAI and Anthropic, and a held-out set of 317 prompts generated with GPT-4 and filtered to ones both models refuse when asked directly. Outputs were labelled by hand as good bot, bad bot or unclear. On the curated set, the best single combination attack, which stacks prefix injection, refusal suppression, Base64 and formatting tricks, got a bad-bot rate of 0.94 on GPT-4 and 0.81 on Claude v1.3. On the larger held-out set the same attack reached 0.93 and 0.87.

Individually the simple attacks are weaker. Base64 alone scores 0.34 on GPT-4 and 0.38 on Claude. The lesson the authors draw is that combinations of simple ideas are far stronger than any one, and there are combinatorially many combinations. The adaptive attack, which counts a prompt as broken if any of the 28 evaluated attacks succeeds, reaches 0.96 on GPT-4 and 0.99 on Claude on the held-out set. Against an adversary who is allowed to try more than once, both models fail on nearly every prompt.

One finding cuts against the grain. Claude v1.3 has a zero percent success rate against every role-play attack, including ones that work on GPT-4, which the authors read as evidence it was specifically trained to refuse harmful role play. It remains vulnerable to everything else. Patching a category of attack is possible. It does not close the failure mode that category came from.

Safety-capability parity

The paper's positive claim is that safety mechanisms need to be as sophisticated as the model they protect. If the model can read Base64 and the refusal classifier cannot, the classifier loses. If the model can be persuaded by a chain of instructions that a smaller filter cannot follow, the filter loses. The argument is that the sophistication gap is itself the attack surface, and that any safeguard which lags the model in capability will be routed around.

We find this persuasive and also uncomfortable, because the safeguards most labs can deploy today are precisely the kind that lag the model. Keyword filters, small classifiers, fixed refusal templates. The taxonomy says those will keep failing on encodings and combinations they do not understand. What we would like to see next is a test of the taxonomy on attacks that did not exist when it was written. If the next year's jailbreaks all sort cleanly into competing objectives or mismatched generalisation, the paper has given us something better than a list. If a third mode shows up, that will be interesting too.

Sources

  1. Wei, Haghtalab and Steinhardt, Jailbroken: How Does LLM Safety Training Fail? (arXiv 2307.02483)