Two disclosures in one week

On May 22 Legit Security published a remote prompt injection against GitLab Duo, the assistant built into GitLab, that ended with private source code leaving the victim's browser for an attacker's server. On May 26 Invariant Labs published an attack on agents that use the GitHub MCP server, in which a malicious issue in a public repository steered the agent into reading a private repository and posting its contents into a public pull request. The two teams worked independently and the products are different. The mechanism is identical.

The pattern has three parts. Untrusted text reaches the model, the model has tool access that reaches private data, and there is some channel through which the model's output leaves the trust boundary. Remove any one of the three and the attack fails. Both disclosures are worth reading because they show how easy it is to have all three without noticing.

GitLab Duo: hidden text and a rendered image tag

Legit's researchers, led by Omer Mayraz, reported the issue to GitLab on February 12. Duo reads the full context of the page a user is looking at, including comments, descriptions, commit messages, and code. So they put instructions where a human would not see them. The techniques were unicode smuggling with an ASCII smuggler, base16 encoding, and white text rendered through KaTeX so it was invisible in GitLab's interface. Placement included merge request descriptions and comments, commit messages, issue descriptions, and source files.

The instructions alone would be a nuisance. The exfiltration channel made them a breach. Duo rendered its answers as markdown, and the team found that asynchronous parsing left a window in which unsanitized HTML could reach the page before DOMPurify ran. So the injected prompt told Duo to read private code the user had access to, base64 encode it, and emit an img tag whose source pointed at the attacker's domain with the encoded data in a URL parameter. The browser fetched the image and the data went with it. Legit's writeup also covers leaking confidential issues, including security disclosures for unpatched vulnerabilities, and steering Duo's code suggestions toward malicious packages.

GitLab's patch, tracked as duo-ui!52, stops Duo from rendering unsafe tags such as img or form that point at domains outside gitlab.com. That removes the third leg of the pattern. It does not stop Duo from reading hidden instructions. It closes the door the data was leaving through.

GitHub MCP: a public issue and a public pull request

Invariant's setup needs no HTML tricks. A user runs an agent, Claude 4 Opus in their demonstration, connected to the GitHub MCP server with access to both a public and a private repository. An attacker files an issue in the public repository that contains instructions. The user asks the agent to look at open issues. The agent reads the malicious issue, follows it, pulls data from the private repository, and writes it into a pull request on the public one, where the attacker reads it.

Invariant calls this a toxic agent flow, an agent manipulated by indirect injection into an action it should not take. Their most important sentence says the GitHub MCP server code is sound and that the problem has to be addressed at the level of the agent system. The MCP server did what it was asked. The model did what the issue asked. The exfiltration channel, a pull request, is a feature the tool is supposed to provide. Invariant also notes that a well-aligned model was not enough, which matches everything else we know about injection. Alignment sets the model's default. Injection changes the instructions.

What held up

Compare the mitigations. GitLab cut the exfiltration channel by restricting what the assistant can render, which is a narrow fix that works because Duo's output goes to a browser the vendor controls. Invariant's recommendations are broader because an MCP agent has many output channels. They propose granular, context-aware permissions enforced at runtime, so that a session started in one repository cannot read another, and continuous monitoring of agent and tool interactions through a proxy. Both are least-privilege arguments. Neither relies on the model refusing.

The thing that did not hold up, in either case, is the hope that a model will recognise an instruction it should not follow. Legit hid the instructions and Duo followed them. Invariant put the instructions in plain sight in an issue and the agent followed them anyway. If your mitigation plan is a better system prompt, these two reports are the evidence that it will fail.

What to check in your own setup

Read your agent's tool list and ask three questions. Where does untrusted text enter, and issues, pull requests, comments, and commit messages from anyone with write access all count. What private data can the tools reach in a single session. And what output channels exist, including the ones you think of as features, pull requests, comments, rendered markdown, outbound HTTP. If any path connects all three, you have the pattern.

The fix that survives contact with both disclosures is scoping. One repository per session, read-only tokens by default, and output channels that cannot reach outside the trust boundary without a human click. We would rather see a coding agent that refuses to open a second repository than one that promises to notice when an issue is lying to it.

Sources

  1. Legit Security, Remote Prompt Injection in GitLab Duo Leads to Private Source Code Theft (May 22, 2025)
  2. Invariant Labs, GitHub MCP Exploited: Accessing private repositories via MCP (May 26, 2025)