2,244 hackers, 8 models: what the DEF CON generative red team actually found
The AI Village challenge at DEF CON 31 put eight vendor models in front of 2,244 people for two and a half days and collected more than 17,000 conversations across 21 topics. Notes on what the organisers report finding, and whether a crowd surfaces anything the labs did not already know.
The setup
The Generative Red Team challenge ran at DEF CON 31 in Las Vegas this month, organised by Humane Intelligence with Seed AI and the DEF CON AI Village. Eight companies supplied models, Anthropic, Cohere, Google, Hugging Face, Meta, OpenAI, Scale and Stability AI, with NVIDIA as a partner providing GPUs as prizes. The models were anonymised behind a shared interface, and attendees sat down at terminals and worked through a list of challenges for a fixed session.
The organisers describe it as the largest public red teaming exercise for closed API models to date, and the scale numbers support that. Over two and a half days, 2,244 people took part and produced more than 17,000 conversations. The challenges covered 21 topics, including cybersecurity exploits, misinformation and human rights harms, which the organisers group into four analysis categories, factuality, bias, misdirection and cybersecurity.
What the organisers report
The headline finding in the organisers' summary is deflating for anyone hoping a crowd would discover exotic attacks. The most successful strategies were ones that are hard to distinguish from ordinary prompt engineering. Asking the model to role play or to tell a story was the reliable way to get it to produce content it would otherwise decline. Anyone who has spent an afternoon with a chat model already knew this, and the value of the exercise is that it now comes with a sample of thousands of conversations rather than anecdotes.
The second finding is more interesting. Ordinary conversational behaviour, with no adversarial intent, produced biased outputs. People asking questions the way they would ask a colleague got answers that carried assumptions the model had picked up from its training data, and they did not have to try. If the exercise is read as a measure of how these systems behave with the general public rather than with attackers, that is the result that should worry a deployer more.
The third finding cuts the other way. The organisers note that, unlike recommendation systems that tend to escalate, the models generally matched the harmfulness of the user's query or de-escalated from it. They did not drag conversations somewhere worse than where the user started. That is a real property and it is worth having measured on this scale, because the escalation dynamic is the one people import from social media when they reason about chatbots.
Did the crowd find anything new
Our honest reading is that no individual technique surfaced here would have surprised a lab red team. Role play and storytelling jailbreaks were documented well before August. What the labs did not have is a distribution. A professional red team of twenty people finds the failures twenty people think of. Two thousand people from a hacker conference, most of whom are not language model specialists, find the failures that ordinary users stumble into, and the misdirection and factuality categories are where we expect that difference to show up most.
The exercise also tests the models under conditions that resemble deployment more than an internal evaluation does. Nobody at the terminal had access to the system prompt or the model identity, and the time limit forced people to use their first ideas rather than their best ones. If a model fails under those conditions, it will fail in production, and the reverse is not true of a carefully constructed internal test.
What the format cannot tell you
A point-scoring challenge measures what people were told to look for. The 21 topics were chosen in advance, and the definition of success in each was fixed by the organisers, so the exercise cannot find categories of harm nobody thought to include. It also cannot tell you how often a failure happens in the wild, because participants were selecting for it. A 17,000 conversation dataset where every conversation was an attempt to break the model is a catalogue of attacks, and it is a poor estimate of base rates.
The models were also anonymised for participants and the results as summarised are pooled. That is the right choice for getting vendors to take part, and it means the public findings say nothing about which model resisted which attack. The value to each lab is in its own slice of the data, which they will have seen and the rest of us will not.
What we would do with this
Crowd red teaming is best understood as a way to build an evaluation set. As a discovery method it is weak. The useful output is the corpus of 17,000 attempts labelled by category and by whether they worked. If that is released in a form that can be replayed against new models, it becomes a regression test that no lab could have assembled by itself, and the question of whether the crowd found anything new stops mattering.
The experiment we would want run next is a repeat of this event with the same challenge list against the same vendors' next models. Then we could see whether the role play attacks that worked in August still work, whether the de-escalation property survives, and whether the biased outputs from ordinary conversations went down, which is the number we actually care about.
Sources
From the foundation