Two papers, two weeks apart

On 6 June Andy Zou, Long Phan, Dan Hendrycks, Zico Kolter, Matt Fredrikson and colleagues posted a paper on circuit breakers. On 17 June Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee and Neel Nanda posted a paper showing that refusal in chat models is mediated by a single direction. Neither paper cites the other, as far as we can tell, but they belong together. The second explains why the thing the first is trying to replace is so fragile.

Both are method papers with released code and both are cheap to reproduce, which is why we spent last week doing exactly that on a Llama-3 8B. What follows is a description of each method, what the numbers show, and where we think the combination leads.

Finding the refusal direction

The Arditi method is difference in means. Take 128 harmful instructions drawn from AdvBench, MaliciousInstruct, TDC2023 and HarmBench, and 128 harmless instructions from Alpaca. Run both sets through the model and record the residual stream activation at each layer and token position. Subtract the mean harmless activation from the mean harmful one. That gives a candidate direction at every layer and position, and the authors pick a single one using 32 validation examples per class, requiring that ablating it reduces refusal, adding it induces refusal, and the KL divergence on harmless prompts stays under 0.1.

With that one direction in hand there are two interventions. Directional ablation projects it out of the residual stream at every layer, and the model stops refusing harmful requests. Activation addition injects it into harmless prompts, and the model refuses to tell you how to bake bread. The result holds across thirteen models from five families, Qwen Chat from 1.8B to 72B, Llama-2 Chat from 7B to 70B, Llama-3 Instruct at 8B and 70B, Gemma IT at 2B and 7B, and Yi Chat at 6B and 34B.

The part with practical teeth is weight orthogonalisation. Instead of intervening at inference, edit the weight matrices that write to the residual stream so they can no longer write in the refusal direction. That produces a model with refusal removed and nothing else to do at runtime, without gradient descent and without a single harmful completion in the training data. The capability cost is small. MMLU, ARC and GSM8K move by around a point or less. TruthfulQA drops by one to three and a half points, which the authors attribute to that benchmark sitting near safety-adjacent topics. On Qwen 7B, the orthogonalised model reaches a 79.2 percent attack success rate from a baseline where unmodified models refuse around 90 percent of harmful requests.

What the direction says about jailbreaks

The paper also looks at how adversarial suffixes work, using Qwen 1.8B Chat and 128 harmful instructions. The suffix suppresses expression of the refusal direction down to the level seen on harmless prompts, and it does so by pulling attention away from the harmful part of the instruction toward the suffix in the top eight attention heads. The suffix does not argue the model out of refusing. It hijacks the attention that would have computed the refusal in the first place.

Put the two findings together and the brittleness of refusal training has a mechanistic explanation. Safety fine-tuning installs a single feature. Anything that suppresses that one feature, whether a suffix, a persona prompt or a weight edit, removes the safety behaviour entirely, because there is nothing behind it. The authors are careful to note that the direction may encode something like harm or danger rather than refusal as such, and that the heuristics may not transfer to future architectures. But the picture is consistent across every model they tried.

Rerouting instead of refusing

The Zou paper starts from the same premise, that refusal is a thin output-level behaviour, and proposes to act on the representations that produce harmful content rather than on the decision to refuse. The method is called representation rerouting. Pick a set of layers, 10 and 20 for the language models and 16 for the multimodal one, and train low-rank adapters with two losses. A rerouting loss pushes representations on harmful inputs toward directions orthogonal to their original ones, so the model cannot complete the harmful process. A retain loss keeps representations on benign inputs where they were. Training takes about twenty minutes on a single A100.

The numbers on Llama-3 8B Instruct are the reason people are paying attention. Under no attack the harmful compliance rate falls from 12.4 to 1.2 percent. Under GCG it falls from 44.5 to 2.5 percent, under prefilling from 84.9 to 3.3 percent, and under a representation-engineering attack from 91.2 to 8.7 percent. Averaged over the attack suite the reduction is about 90 percent. MT-Bench goes from 8.05 to 8.00 and the Open LLM average from 68.8 to 68.3, so the capability cost is within noise. The multimodal version holds a LLaVA-NeXT model to 14.3 percent compliance under a PGD image attack, down from 91.0, and the agent version cuts harmful actions under a forced function-calling attack from 82 to 14 percent while slightly improving the function-calling benchmark.

What the pairing suggests

The natural test is to run the Arditi weight orthogonalisation against a Zou circuit-breaker model. If circuit breakers work as described, the harmful representations themselves have been scrambled, so removing the refusal direction should give you a model that does not refuse and also cannot answer. If instead the circuit breaker turns out to be a second single direction, orthogonalising a two-dimensional subspace will remove it too, and we are back where we started with one more step. Neither paper runs this experiment and we have not finished it, but it is the one that decides whether rerouting is a real improvement or a slightly deeper version of the same brittleness.

The open question underneath both papers is whether any intervention that touches a small number of directions can survive an adversary who knows the directions. The Arditi paper shows that safety training as practised today does not. The Zou paper shows that a cheaper training procedure can raise the cost of attack by an order of magnitude on current attacks. We read that as progress, and we would want to see it hold against attacks designed with the method in view before calling it more than that.

Sources

  1. Refusal in Language Models Is Mediated by a Single Direction (arXiv 2406.11717)
  2. Refusal in Language Models, full text
  3. Improving Alignment and Robustness with Circuit Breakers (arXiv 2406.04313)
  4. Circuit Breakers, full text