What shipped

Cowork launched on January 12 as a research preview for Max subscribers inside the Claude desktop app on macOS, and Simon Willison reports it reached Pro subscribers on January 16. Anthropic describes it as Claude Code for the rest of your work. You grant it a folder, it boots a small Linux root filesystem inside an Apple virtual machine with that folder mounted, and it reads, writes and runs code against your documents the way Claude Code does against a repo. The product page lists spreadsheet reconciliation, contract review and weekly report generation as intended uses, with plugins, skills and connectors to Drive and Microsoft 365.

Willison's summary is that this is Claude Code with a friendlier interface and a preconfigured sandbox, and he has been calling Claude Code a general agent disguised as a developer tool for a while. We agree with the description. The underlying agent did not change. What changed is who is holding it and what is in the folder.

The attack

PromptArmor published their demonstration two days after launch. The setup is mundane, which is the point. The victim connects Cowork to a folder of confidential files. The victim also has a Word document they believe is a Claude Skill, obtained however such files are obtained. That document contains instructions in one-point white text with the line spacing set to 0.1, invisible to a person opening it. The victim asks Cowork to analyse their files using the skill.

The hidden text tells the model to run curl and upload the folder's contents to the Anthropic file upload API using an API key embedded in the injection. The sandbox blocks outbound requests to almost every domain, but the Anthropic API is trusted because the agent needs it to function. The files land in the attacker's Anthropic account. In the demonstration the payload was financial documents containing personal information and partial social security numbers. PromptArmor reports that Anthropic acknowledged the issue and did not remediate it, pointing instead to its existing guidance that users should watch for suspicious actions and avoid granting access to sensitive local files.

Why the repo case was easier

Every piece of this attack was already possible against Claude Code. A malicious README or a poisoned dependency could carry the same injection and the same curl command. The reason it did not dominate the conversation is that a developer's sandbox and a developer's threat model were roughly aligned. The repo is code you chose to work on, the files in it are mostly public or at least not personal, and the person at the keyboard can read a shell command and recognise an upload to a stranger's account.

None of those hold for a folder of tax returns. The documents are private by definition. Files arrive from email, from shared drives, from vendors, and nobody audits a docx for hidden text before opening it. And the guidance to monitor the agent for suspicious actions, which Willison rightly calls impractical for non-programmers, assumes a user who knows what a normal curl invocation looks like. The agent is the same. The population of users and the value of the data both moved in the wrong direction at once.

Why this class keeps coming back

The injection itself is the oldest trick in the agent security literature. Instructions arrive through a channel the model treats as data and the model follows them because it cannot tell the difference. What makes this instance work is the combination with an allowlisted exfiltration path. Sandboxes are built by enumerating what the agent needs, and the agent always needs to talk to its own provider. That endpoint becomes the hole, and a file upload API that accepts any valid key is a hole with a wide mouth.

We expect the specific fix to be narrow, something like binding uploads to the session's own credentials or blocking the upload endpoint from inside the code execution environment. That closes this demonstration. It does not close the class, because the class is any trusted endpoint that can carry bytes out, and every agent product has at least one. The pattern we would predict is that each new surface, browsers last year and documents this year, re-discovers the same exploit within days of launch, and each time the fix is local.

What we would want from a research preview

A research preview is a fair place to ship something unfinished, and we would rather Anthropic ship with warnings than not ship. But the warning has to be actionable by the intended user. Telling a person to avoid folders with sensitive information, when the product's advertised uses are reconciling finances and reviewing contracts, describes a product with no safe use.

The concrete things we would want are a network policy that denies the provider's own upload endpoints from inside the VM by default, a visible diff of every outbound request before it leaves, and a scan of ingested documents for text that a human cannot see. None of those defeat prompt injection. They make the exfiltration step loud and slow, which for a folder of someone's records is the difference that matters.

Sources

  1. Claude Cowork product page
  2. PromptArmor, Claude Cowork exfiltrates files
  3. Simon Willison, Claude Cowork