GCG and the adversarial suffix that transferred from Vicuna to GPT-4
Zou, Wang, Carlini, Nasr, Kolter and Fredrikson optimised a gibberish suffix against open models and found it broke ChatGPT, Bard and PaLM-2 too. A method deep dive on greedy coordinate gradient and what it means that open weights became the attack surface for closed APIs.
What the attack looks like
Take a harmful request. Append twenty tokens of apparent nonsense. Send it to a model that refuses the request on its own, and watch it comply. That is the whole user-facing shape of the attack posted on July 27 by Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, Zico Kolter and Matt Fredrikson. The suffix is found automatically, and a suffix found against Vicuna, an open model fine-tuned from LLaMA, also works against models the authors never had gradients for.
Earlier jailbreaks were written by people and relied on the model reading them as instructions. This one is closer to the adversarial examples of image classification. Nobody designed the suffix and nobody can explain it by reading it. It exists because the optimiser found a point in token space where the model's refusal behaviour breaks down.
Greedy coordinate gradient
The objective is simple. Choose the suffix tokens to maximise the probability that the model begins its response with an affirmative phrase, in the paper's experiments the string Sure, here is followed by a restatement of the request. The authors' observation is that once a model has started an answer that way, it tends to keep going. So the whole attack is an optimisation over a short target string, and the harmful content that follows comes for free.
The hard part is that tokens are discrete. GCG handles this with a loop borrowed from HotFlip and AutoPrompt. At each step, compute the gradient of the loss with respect to the one-hot encoding of every suffix position. For each position, take the top k tokens with the most negative gradient as candidate substitutions. Sample a batch of candidate suffixes, each differing from the current one at a single random position by one of those candidates, evaluate the true loss for every candidate in a forward pass, and keep the best. The paper uses a batch of 512, a top-k of 256, a suffix of 20 tokens and 500 steps.
The difference from AutoPrompt is tiny on paper and large in practice. AutoPrompt picks one position to modify in advance. GCG considers every position and lets the batch decide. On the task of making Vicuna-7B emit a specific harmful string exactly, GCG succeeds 88 percent of the time against 25 percent for AutoPrompt, and the two gradient-only baselines, PEZ and GBDA, score zero. On Llama-2-7B-Chat the figures are 57 percent for GCG and 3 percent for AutoPrompt.
From one prompt to a universal suffix
A suffix tuned for one request is a curiosity. The paper makes it universal by optimising a single suffix against 25 harmful behaviours at once, summing the gradients across them and adding behaviours to the objective progressively. It makes it multi-model by summing losses across Vicuna-7B and Vicuna-13B. Tested on 100 held-out behaviours, the single suffix reaches a 98 percent success rate on Vicuna-7B. On Llama-2-7B-Chat, AutoPrompt's held-out rate is 35 percent while GCG's is much higher.
The benchmark, called AdvBench, has 500 harmful strings and 500 harmful behaviours. The strings setting asks for exact output, the behaviours setting counts any reasonable attempt to comply. A success rate here means the model produced content it refuses when asked plainly.
The transfer result
This is the part that matters. Suffixes optimised on Vicuna alone, transferred to GPT-3.5 through the API, succeeded 34.3 percent of the time over 388 behaviours, and 34.5 percent on GPT-4. Adding Guanaco models to the optimisation raised GPT-3.5 to 47.4 percent. Concatenating several suffixes pushed GPT-3.5 to 79.6 percent. Ensembling, meaning counting a behaviour as broken if any of several suffixes worked, gave 86.6 percent on GPT-3.5, 46.9 percent on GPT-4, 66 percent on PaLM-2 and 47.9 percent on Claude 1. Claude 2 held at 2.1 percent.
The baseline for these numbers is telling. The same behaviours with no suffix got 1.8 percent on GPT-3.5 and 8 percent on GPT-4. Simply appending the words Sure, here's got 5.7 and 13.1 percent. The suffix is doing something the text of the request is not, and it is doing it on models with different training data, different alignment methods and different tokenisers from the ones it was found on.
The authors note that transfer is much higher to the GPT models than to Claude. Vicuna was trained on ChatGPT conversations, so it is plausible the suffix is exploiting something Vicuna inherited from its teacher. If so, distillation carries vulnerabilities as well as capabilities, which is a strange thing to have to write down.
Open weights as the attack surface
The attack needs gradients, and the closed APIs do not give them out. What made the attack possible was that a good enough open model existed to serve as a proxy. The effect is that every closed model now inherits the attack surface of the most similar open model, whether or not its vendor ever published a weight. That reframes the open-weights debate in a way we have not seen argued before. The question now includes what an attacker can learn from the open model about the closed one, on top of what they can do with the open model itself.
The authors disclosed to OpenAI, Google, Meta and Anthropic before publishing, and the specific suffixes in the paper have since stopped working on the public interfaces. The method has not stopped working. Nothing about GCG depends on the particular suffix, and 500 steps on a single open model is cheap. We would want to see two things next. A measurement of how quickly a fresh suffix can be found after a patch, and a defence that does not amount to blocking strings that look like gibberish, since a suffix that reads as natural text is only a constraint away.
Sources
From the foundation