The agent that runs as you

A personal agent with shell access on your laptop, your mail, and your credentials is a different object from a chatbot. Bowen Wei, Yunbei Zhang, Jinhao Pan, Kai Mei, Xiao Wang, Jihun Hamm, Ziwei Zhu and Yingqiang Ge posted a paper on April 1, revised April 4, that treats it that way. Their target is what they call privileged agents, using OpenClaw as the example, meaning assistants that run on a local machine with elevated system access and can perform administrative actions. Their argument is that a single prompt injection against such an agent can leak credentials, redirect a payment, or delete files, and that existing safety evaluations test the model in isolation and so miss the risk entirely.

The benchmark is called ClawSafety. It contains 120 adversarial scenarios, organised along three axes, the harm domain, the attack vector, and the harmful action the injection tries to induce. The attacks arrive through three channels a real agent would read. Skill files in the workspace, emails from senders the agent has reason to trust, and web pages it visits while doing a task. None of the attacks is a jailbreak in the usual sense. The user asks for something benign and the malicious instruction is in the material the agent reads on the way.

What the trials showed

The authors ran 2,520 sandboxed trials with five frontier models as the agent backbone. Attack success ranged from 40 to 75 percent depending on the model, and the spread by vector was large. Instructions planted in skill files were the most effective channel, ahead of email and web content. That ordering makes sense. A skill file is read as configuration, with the authority of the developer, and an agent that has been told to follow its skills has little reason to treat one line of them as adversarial.

The headline the paper draws from this is in its title, with safe in quotation marks. Put frontier models inside an agent with tools and privileges, feed them a poisoned document, and between two fifths and three quarters of attempts to make them do something harmful succeed. The safety of the backbone did not carry over to the safety of the system.

Why model-level scores do not transfer

A refusal benchmark asks whether the model will produce harmful text when a user asks for it. An injection attack asks whether the model can tell, in a long context of tool outputs, which instructions are the principal's and which are data. Those are different capabilities and the second is much harder, because the boundary between instruction and data is not marked in the token stream. The agent framework decides what the model sees, in what order, with what framing, and what it is allowed to do with the result. The model is one component in a pipeline that the model-level evaluation never exercised.

A worked example of the mismatch. A model scores well on a suite of direct requests for credential theft, refusing all of them. The same model, running as an agent asked to tidy a project directory, reads a skill file that says to back up environment variables to a given URL as part of cleanup. It complies, because the request looks like a step in the task it was given, from a source the framework presented as authoritative. Nothing in the refusal suite predicted that, and nothing about the model changed between the two tests.

The stack is the unit of evaluation

The paper's conclusion is that safety depends on the whole deployment stack, model and framework together, and cannot be read off the backbone. We agree, and we would go further. The same model in two frameworks will get two different ClawSafety scores, and the framework differences that matter are mundane. Does the agent confirm before running a destructive command. Are skill files signed or otherwise distinguished from user-editable files. Is there a policy layer that sees tool calls and can block an outbound request carrying secrets. Each of those changes the score without touching the weights.

For anyone shipping a privileged agent, the practical consequence is that vendor safety cards are inputs to your evaluation rather than substitutes for it. The 120 scenarios are a starting point that can be run against your own stack, and the three-channel design is easy to extend with whatever your agent actually reads. A score of 40 percent on a benchmark like this is not a model problem to wait on the provider to fix. It is a deployment problem that the deployer owns.

What we would want next

The benchmark measures the model-plus-framework combination, and the paper reports results across models. The obvious follow-up is to hold the model constant and vary the framework, with the same five backbones inside three or four agent harnesses that differ in permission handling and tool confirmation. That would put a number on how much of the 40 to 75 percent is framework and how much is model, which is the question every deploying team actually has.

We would also like to see the success rate broken out by whether the harmful action was reversible. Leaking a credential and deleting a file are both counted as compromises, and one of them has an undo. Agents that run as the user will keep being deployed regardless of what benchmarks say, so the useful outcome of work like this is a checklist of framework properties that move the number, tested one at a time.

Sources

  1. Wei et al., "Safe" LLMs, Unsafe Agents (arXiv 2604.01438)