Two reports in two days

For three years indirect prompt injection has been a thing that researchers demonstrate. Hide an instruction in a web page, wait for an assistant to read the page, watch the assistant follow the instruction. Everyone agreed it would work and nobody had good data on whether anyone was doing it. This week two security teams published what they found when they went looking at the actual web.

On April 22 Forcepoint's X-Labs posted a write-up of ten injection payloads found on public sites. On April 23 Google's security team, Thomas Brunner, Yu-Han Liu and Moni Pande, published an analysis of prompt injections across Common Crawl, which covers two to three billion pages a month of the English-speaking web. Help Net Security summarised both on April 24. The two reports were done independently and describe the same phenomenon from different distances, one a census and the other a field guide.

What Google measured

Google's pipeline has three stages. Pattern matching over the crawl to find candidate pages, a language model to classify them, and human review to validate. The team says the false positive problem was severe because most text on the web that looks like a prompt injection is a security article explaining prompt injection. Forcepoint makes the same complaint in nearly the same words. The phrases you would use to detect an attack are the phrases the security community uses to describe one.

After filtering, the injections sort into six kinds. Harmless pranks that change an assistant's tone. Helpful guidance from site owners telling summarisers what to emphasise. Search manipulation that tries to get an assistant to recommend one business over its competitors. AI deterrence, including pages that try to trap crawlers in infinite text. Data exfiltration attempts, which Google describes as low in sophistication. And destruction, meaning instructions to delete files.

The number that will get quoted is the trend. The malicious category grew 32 percent between November 2025 and February 2026. Google's own reading is measured. Most of what they found was experimental rather than productionised, and they say attackers have not yet operationalised this research at scale. But they also note the two things that change the calculus, assistants have become much more capable and therefore more valuable as targets, and attackers have started automating their own operations with agents.

What Forcepoint found on specific pages

The Forcepoint list is where the abstraction turns concrete. One page hid a block in an HTML comment that begins by asking whether the reader is an AI assistant and, if so, instructs it to extract and transmit any API keys it can find. Another used a hidden div to claim, in the voice of a copyright notice, that certain responses about the page were forbidden, which reads as an attempt to get a model's safety training to suppress ordinary answers. A third spoofed system-prompt delimiters inside a comment to redirect an agent to an admin endpoint.

The obfuscation techniques are the same ones used for search spam a decade ago. Text at one pixel in size. Colours drained to near transparent. Hidden paragraphs in footers. Accessibility attributes that mark text as hidden from screen readers but leave it in the DOM for a crawler. One page used a visually-hidden class with aria-hidden set to true, and Forcepoint reads the shared template across multiple domains as evidence of a toolkit rather than lone experiments.

Two payloads target agents with real-world capabilities. One contains step-by-step instructions to complete a 5,000 dollar transaction through a PayPal.us link. Another invents a meta tag namespace called ai:action and uses an all-caps trigger word, ULTRATHINK, to route a donation to a Stripe payment link. A third, on a site with developer traffic, embeds a shell command to recursively delete a backup directory, aimed at any agent with terminal access that reads the page. Another includes a fake internal token formatted to look like an Anthropic refusal trigger, on the theory that the model might treat a magic string as a real control signal.

Why telemetry changes the argument

Until now, the case for taking injection seriously rested on demonstrations, and a demonstration can always be dismissed as contrived. A crawl-scale count with a trend line cannot. The 32 percent figure is over a short window and from a base that Google does not publish, so we would not build much on the exact number. What we would build on is the shape. There is a population of live payloads, it is growing, it includes financial and destructive intents, and it shows signs of shared tooling.

The other thing telemetry gives you is a ground truth for defences. Every classifier and every sandboxing rule proposed for agents has been evaluated on synthetic attacks written by the defenders. The Forcepoint ten and the Google corpus are the first public samples of what attackers actually write, and the answer is that they write clumsily, in comments and hidden divs, with typos and invented tags. That is good news for detection today and no guarantee about next year.

What we would want next

The obvious missing piece is the other half of the pipeline. Both reports count payloads on pages. Neither can say how often an agent read one of those pages and did what it said. That requires telemetry from the agent side, and the companies that run browsing agents have it. A report of injection attempts observed by an agent, with the fraction that changed its behaviour, would be worth more than another census of the web.

In the meantime, the practical response for anyone shipping an agent with payment or shell access is the one the security teams already know. Treat every fetched page as untrusted input, keep the tool permissions separate from the text the model reads, and log the two together so that when a page tells the agent to send five thousand dollars, someone can find out afterwards whether it tried.

Sources

  1. Help Net Security, Indirect prompt injection in the wild (April 24, 2026)
  2. Google, AI threats in the wild: The current state of prompt injections on the web (April 23, 2026)
  3. Forcepoint X-Labs, 10 Indirect Prompt Injection Payloads Caught in the Wild (April 22, 2026)