What was shown in Las Vegas

Artem Chaikin, a security engineer at Brave, gave a session at Black Hat USA 2026 titled Attacking and Defending AI Browsers. The list of targets was short and familiar: Opera's AI browser, Perplexity's Comet, and OpenAI's ChatGPT Atlas. Every one of them fell to indirect prompt injection, meaning the attacker never touched the user's prompt. They put text on a page and waited for the agent to read it.

The techniques were not exotic. Opera's guardrails were bypassed with instructions hidden behind HTML. Comet was steered by nearly invisible text laid over an image and by instructions tucked inside Reddit comments behind spoiler tags. Atlas was fooled by content that mimicked its own trusted content tags, and exfiltrated data slipped past its scanning when it was packed into URL fragments. None of this needed a zero-day. It needed a page the agent would visit and a model that treats what it reads as something to act on.

The pay to live detail

The finding that stuck with us was about Atlas and usage caps. Free users who hit their limit were downgraded to a less capable model, and the less capable model was also less resistant to injection. Chaikin put it as a question the audience was clearly already asking: is this some kind of pay to live scenario? Security that depends on which model tier you happen to be routed to is security that changes under the user's feet without telling them.

This matters beyond Atlas because routing between models on cost is now standard practice. If injection resistance varies by model, and it does, then every product that silently swaps models has a security posture that varies by time of day and account status. We have not seen a vendor publish injection resistance per tier. That would be a useful number to demand.

Brave had already said the quiet part in June

Two months before Black Hat, Brave published a research post on indirect injection that is worth reading alongside the talk. Their two case studies were not browsers. One was Mozilla's Tabstack, a cloud hosted web execution API, and the other was Cotypist, an on-device macOS autocomplete tool. Tabstack was hijacked with white on white and zero width text on a page, after which the agent abandoned its task, navigated to an attacker domain, filled a form with the conversation history and submitted it. Cotypist was steered by instructions in a local document, which shaped its suggestions and risked surfacing credentials.

The pairing was the point. One system is cloud, one is local, and both broke the same way. Brave's framing was that the question to ask of any system is whether it composes trusted instructions with untrusted content in a shared context window, and that whether it uses a cloud API is beside the point. Tabstack confirmed the bug on May 14 and a fix on June 1. Cotypist confirmed on June 2. The post ended with a sentence vendors usually avoid: indirect prompt injection cannot be fully solved within the current LLM architecture.

What defending looks like when you cannot win

Brave's own mitigations, as Chaikin described them, are all about limiting blast radius rather than stopping the injection. Agentic browsing runs in a separate browser profile with personal accounts logged out by default, so a hijacked agent has fewer sessions to abuse. There is a minimum model threshold, Claude Haiku 4.5 or higher, which is a direct response to the tier downgrade problem. And there is a secondary sentinel model that checks whether the agent's actions still line up with what the user asked for in language.

The sentinel is the most interesting and the least convincing. A second model reading the same untrusted content can be injected too, and alignment checking in natural language inherits every ambiguity of natural language. It raises the cost of an attack. It does not close the class. Brave's June post said as much when it pointed to structural separation, least privilege and information flow control as the real work, and those are system design changes, not prompt changes.

What we would want measured next

Chaikin's closing line was that we just have to deal with the uncertainty here. We think that is the honest position, and it is a strange one for a security talk to end on. What would make the uncertainty smaller is a public, repeatable injection benchmark run against shipping browsers, with results broken out by model tier and by the channel the injection arrived through, whether page HTML, image overlay, comment thread or URL. Every vendor in the talk fixed the specific bug reported to them. None of them can say what fraction of the attack surface that bug represented.

The other thing we would try is treating the profile isolation idea as a measurement rather than a mitigation. Run the same injection suite with the agent logged in to everything and with it logged in to nothing, and publish the difference in realised harm. If the gap is large, that argues for making the logged out default a hard requirement across products. If it is small, the problem is elsewhere and the industry should stop selling isolation as the answer.

Sources

  1. Dark Reading, No perfect fix for AI browser prompt injection flaws (Aug 5, 2026)
  2. Brave, Indirect prompt injection remains a fundamental security challenge for AI (Jun 8, 2026)