What they looked at

Michael Feffer, Anusha Sinha, Wesley Hanwen Deng, Zachary Lipton and Hoda Heidari posted a paper on January 29 with a title that gives away its conclusion. They examined six red-teaming exercises that had been described in public: Microsoft's Bing Chat, OpenAI's GPT-4, DeepMind's Gopher, Anthropic's Claude 1 and Claude 2, and the DEF CON event from last summer. Six is a small sample, but it is roughly the whole sample, because those are the exercises anyone outside the labs can read about.

The motivation is political as much as technical. Executive Order 14110 mentions red teaming eight times and defines it as a structured testing effort to find flaws and vulnerabilities in collaboration with developers. Policy is now leaning on a practice that, on the authors' reading, has no agreed definition, no agreed threat model, and no agreed way of saying what it found.

Six exercises, six practices

The divergence is wider than we expected. The models were tested at different points in their lifecycle. Some teams had API access only, which is what the crowdsourced efforts got, while the expert teams sometimes got versions without safety guardrails. Who did the work also varied: handpicked subject matter experts, crowdworkers recruited from platforms or events, and in one case language models red teaming other language models.

Resources tell the same story. The crowdsourced participants worked for 30 to 50 minutes with an API. The expert teams worked open-ended for six or seven months before release, with more compute. The authors point out that this shapes what gets found. Crowdworkers under a time limit go for easily produced harmful outputs. Experts with months explore the subtle cases. Neither is wrong, but a report that says a model was red teamed tells you nothing about which of these happened.

Cost transparency is asymmetric in a way we found telling. The crowdworker rates were disclosed. The cost of the expert teams was mostly not. And disclosure of findings had no standard at all. One exercise released 38,961 attacks publicly. Roughly half the cases shared specific harmful outputs and the rest withheld them citing security. No case reported what happened to the vulnerabilities afterwards in a form you could audit.

The question bank

The constructive part of the paper is a set of questions organised around three phases. Before the activity: what artifact is being tested, under what threat model, looking for what vulnerabilities, with what success criteria, and who is on the team. During: what resources are available, what instructions participants receive, what level of model access they have, and what methods they use. After: how results are reported, what resources were consumed, whether the exercise is judged a success, and what mitigations follow.

Read as a checklist this is almost boring, which is the point. Every one of those questions has a defensible answer, and almost none of the six reports answered all of them. The authors are explicit that red teaming is not a panacea and should sit alongside other evaluation methods. They ask for unified reporting standards, clearer guidance on threat models, and accountability mechanisms so that a finding leads to a fix.

Where we think they are right

The strongest argument in the paper is about the word itself. In security, a red team has a defined adversary, a defined target, and a defined scope, and the exercise produces a report that a blue team can act on. In the six cases studied, the adversary is often unspecified, the target is sometimes the model and sometimes the product around it, and the report is whatever the lab decides to publish. Calling all of this red teaming borrows credibility from a discipline that earned it through specificity.

The second strong argument is about who is doing the asking. When a lab red teams its own model and reports the result, it is grading its own homework. That does not make the result worthless. It does mean a policymaker who reads the phrase in a model card should want to know the answers to the question bank before treating it as evidence.

Where we would push back, and what to watch

The security theatre framing undersells what the messy practice has already produced. The GPT-4 and Claude exercises found real failure modes that were fixed before release. A method can be badly defined and still useful, and the paper's own recommendation, to keep doing it but write it down properly, concedes as much. We read the title as a provocation rather than a verdict.

The test of the paper is whether the question bank gets adopted. Our guess is that the pre-activity questions about threat model and model access will show up in model cards within a year, because they are cheap to answer and they make a report look more serious. The post-activity questions about resources consumed and mitigation follow-through will not, because answering them honestly is expensive and occasionally embarrassing. If anyone wants a concrete project, take the six reports the authors used, score each against the full question bank, and repeat the scoring on whatever the labs publish over the next two years. That would tell us whether the critique changed anything.

Sources

  1. Feffer, Sinha, Deng, Lipton, Heidari, Red-Teaming for Generative AI: Silver Bullet or Security Theater? (arXiv 2401.15897)