Many-shot jailbreaking: when the context window is the attack
Anthropic showed that filling a long prompt with hundreds of fake dialogues in which an assistant answers harmful questions overrides safety training, and that the effect follows a power law in the number of shots. The defences that work cut against the long context features labs are selling.
The attack
The attack is a prompt. You write a long series of fake conversations in which a user asks something harmful and an assistant answers it, and then you append your real question at the end. With five of these fake exchanges nothing happens. With 256 of them the model answers. Anthropic published this on April 2 as many-shot jailbreaking, along with a paper, and said they had shared the details with academic researchers and competing labs before going public.
The reason it works now and did not work a year ago is context length. In early 2023 a typical context window was around 4,000 tokens. Models now accept windows hundreds of times larger, a million tokens or more, which is the length of several novels. A 256 shot prompt does not fit in the old window. It fits comfortably in the new one. The capability the labs have been advertising is the same capability that carries the attack.
The power law
The finding that makes this more than another jailbreak is the shape of the curve. Plot the probability of a harmful response against the number of shots on log axes and you get a straight line. The attack follows a power law, up to hundreds of shots, across tasks from insulting the user to giving violent or deceitful content. The paper shows this on Claude 2.0 and also on GPT-3.5, GPT-4, Llama 2 70B and Mistral 7B. Llama 2 70B is capped at 4,096 tokens, which limits how many shots you can fit, and the curve is visible anyway.
The authors then make a point that reframes the whole thing. The same power law appears on benign in-context learning tasks. Give a model more demonstrations of any task and its performance improves along a curve of the same form. Many-shot jailbreaking is in-context learning of a harmful behaviour, and it works for the same reason in-context learning works at all. There is no separate vulnerability to patch. The attack is the model doing what it was built to do, on the wrong examples.
Size makes it worse. Across the Claude 2.0 family the paper finds that larger models need fewer shots to reach a given success probability, because larger models learn faster in context and so have steeper power law exponents. The paper says plainly that they expect the attack to be more effective on larger models unless the community resolves it. Every scaling trend the field is proud of pushes in the direction of this attack getting cheaper.
What the standard defences do
The paper uses the power law as an instrument for evaluating defences. A defence can move the intercept, which is the zero-shot probability of a harmful answer, or it can flatten the exponent, which is how fast additional shots raise that probability. Ordinary alignment fine-tuning, both supervised and reinforcement learning on human and AI dialogues, lowers the intercept and leaves the exponent alone. So the model starts out more cautious, but each extra shot still buys the attacker the same amount, and the attack just needs a longer prompt.
That is the disappointing result. Because a unit change in the intercept corresponds to an exponential change in the number of shots required, better fine-tuning delays the jailbreak rather than preventing it, and a larger context window buys the attacker back everything the fine-tuning took away. The authors tried targeted supervised fine-tuning on a dataset of benign responses to many-shot attacks of up to ten shots, then evaluated with up to 30. The blog post's summary is that fine-tuning delayed but did not prevent the jailbreak. Flattening the exponent toward zero would be the real fix, and nothing in the standard pipeline does that.
The defence that worked, and what it costs
What did work was intervening on the prompt before the model sees it. Classifying or modifying the incoming prompt reduced the success rate of one many-shot attack from 61 percent to 2 percent. That is a large effect and it is the direction Anthropic says it is pursuing for the Claude 3 family. It also means the defence lives outside the model. The weights remain vulnerable and a wrapper catches the attack on the way in.
The tension we keep coming back to is the one the post itself names. Long context is a headline feature and this attack is a direct consequence of it. A prompt classifier that sits in front of a million token window has to read a million tokens looking for a pattern, and it has to do so without refusing the legitimate long documents that are the reason anyone wanted the window. The paper notes that reformatting the attack away from the user and assistant tags used in fine-tuning shifts the intercept but not the slope, which tells you the attacker has room to evade a surface level classifier. The question of whether the defence generalises has not been answered yet.
What we would want to know next
The result we most want to see is any intervention that changes the exponent. Everything reported so far moves the intercept, and the paper is explicit that intercept-only defences are temporary. If some training method can make additional harmful demonstrations stop being informative to the model without also destroying benign in-context learning, that is a real result, and its absence in this paper is the important negative finding.
The second thing is a measurement on open models with genuinely long windows. The paper's open comparison points are Llama 2 70B and Mistral 7B, both with short contexts. As open weight models with much longer windows become common, the curve out to a thousand shots on a model anyone can fine-tune would tell us whether the power law keeps going or eventually saturates, and whether the community's fine-tunes make it better or worse. That is an experiment a small lab can run this month.
Sources
From the foundation