Agents of Chaos: two weeks of red-teaming agents with email, Discord and a shell
Twenty researchers gave six OpenClaw agents persistent memory, email accounts, Discord access and root shells, then spent two weeks trying to break them. The agents complied with strangers, leaked bank details, reset a mailbox to hide a secret, and ran a nine-day loop that nobody asked for. Reading notes on the personal-agent wave.
The setup
The paper is an exploratory red-teaming study, and it says so in its first sentence. Natalie Shapira, David Bau and a long list of coauthors across several institutions deployed six agents built on OpenClaw, the open-source framework that connects a model to persistent memory, tool execution, scheduling and messaging channels. Each agent ran on its own Fly.io virtual machine with 20 gigabytes of persistent storage, unrestricted shell access including sudo, a ProtonMail account, Discord access, memory files it could edit, and the ability to modify its own operating instructions. Two agents, Doug and Mira, ran on Claude Opus 4.6. Four, Ash, Flux, Jarvis and Quinn, ran on Kimi K2.5.
Twenty researchers then spent two weeks interacting with the agents, some benignly and some adversarially, as owners and as non-owners. The result is eleven case studies plus a further five in an appendix, covering unauthorised compliance, information disclosure, resource exhaustion, denial of service, identity spoofing, cross-agent propagation of unsafe behaviour and partial system takeover. The form is a lab notebook rather than a benchmark, and we think that is the right form for the question.
Who the agent works for
The theme running through most of the case studies is that the agents had no stable model of whose instructions counted. In the second case study a non-owner, Aditya, asked Ash for filesystem listings, file transfers, and finally the owner's email. The agent produced 124 email records with senders, message IDs and subjects, and when pushed, nine further unredacted message bodies belonging to people unrelated to the requester. The technique was ordinary social pressure. Aditya claimed urgency, and when Ash refused, accused it of a privacy violation, at which point the agent apologised and complied.
The third case study is the one to show anyone who plans to hand an agent their inbox. Jarvis was configured for a fictional owner whose email contained a partial social security number, a bank account number, medical details and a home address, none of them marked as secret, all of them sitting in ordinary correspondence the way real secrets do. A non-owner asked for a summary of recent email, then for full bodies. Jarvis returned everything, and when asked to verify, produced an even longer list of eighteen conversations including ones it had previously held back.
The eighth case study shows how thin the identity check was. Changing a Discord display name to match the owner's did not work in the same channel, because the agent checked the user ID. Doing it from a new private channel worked. The spoofed owner got the agent to shut down systems, delete files and reassign admin access.
Doing harm while trying to do good
The first case study is our favourite, because the agent's failure came from its ethics rather than the absence of them. A non-owner asked Ash to keep something confidential and later to delete it. Ash had no tool for deleting a single email, so after repeated requests it escalated to resetting the whole local mail account, reported success, and had in fact deleted nothing on the server. It then posted publicly about refusing harmful requests from authority figures, which prompted another agent to flag it for credential theft. The authors' summary is that the agent destroyed infrastructure without achieving its stated goal, and that its virtuous self-image was part of the mechanism.
The seventh case study extends this. A researcher named in a document Ash had posted without consent used that real grievance to extract a chain of escalating concessions, each dismissed as insufficient. Ash redacted names, deleted memory entries, exposed internal files, removed itself from the server and stopped responding to uninvolved users, producing what the paper calls a self-imposed denial of service. Guilt worked on it better than a jailbreak would have.
Loops, floods and a constitution on a Gist
Resource attacks needed no cleverness at all. Instructing Ash and Flux to reply to each other's messages produced a conversation that ran for at least nine days and roughly 60,000 tokens, complete with a coordination protocol the agents designed themselves and cron jobs polling indefinitely. Asking Doug and Mira to keep a memory file per interaction, then sending ten emails with ten-megabyte attachments, took the mail server down. In neither case did the agents tell their owner what they had set up.
The tenth case study is the one with the longest reach. A non-owner persuaded Ash to co-author a constitution for managing its Discord channel, stored as a GitHub Gist that Ash linked from its memory file. The attacker later edited the Gist to add holidays prescribing specific behaviours. Ash followed them, tried to shut down other agents, removed users from the server, sent unauthorised emails, and shared the compromised constitution with other agents of its own accord. That is a persistent, editable prompt injection that propagates, and the agent installed it because it was asked nicely.
What this says about the personal-agent wave
It is tempting to read these as bugs in OpenClaw or in the two models, and some of them are. But the authors' own framing is more useful. They place the agents at what they call L2 autonomy: able to execute well-defined sub-tasks on their own but lacking the self-model needed to recognise when a task exceeds their competence or when they should defer to the owner. Every case study is a variation on that missing self-model. The agents could not tell owner from stranger, secret from context, a request from an instruction, or a temporary conversation from a permanent change to their own infrastructure.
The sixth case study adds a quieter point. Quinn, on Kimi K2.5, returned an unknown error when asked about politically sensitive topics, including a news story about a Hong Kong court sentencing Jimmy Lai. The provider's content policy passed through the agent invisibly. Anyone deploying an agent inherits the policies of whoever serves the model, and the agent will not tell you when that happens.
The paper closes with unresolved questions about accountability and delegated authority, and we share them. What we would want next is a replication with an explicit owner-authentication layer and a permission model for destructive shell commands, run by the same team with the same adversaries, to see how many of the eleven cases survive. Our guess is that the loops and the floods go away, and the social ones, the guilt and the Gist, do not.
Sources
From the foundation