What Cloudflare says it saw

On August 4 Cloudflare published an accusation with a test attached. Its engineers registered brand new domains that had never been indexed, put a robots.txt on each one disallowing all crawlers, and added firewall rules blocking Perplexity's declared user agents. Then they asked Perplexity questions about those domains. It answered with detailed information about the exact content of the pages.

The traffic they traced back had two shapes. The declared crawler, identifying as Perplexity-User, was making 20 to 25 million requests a day. Alongside it was a second stream of 3 to 6 million requests a day that presented a generic Chrome on macOS user agent, came from IP addresses outside Perplexity's published range, and switched autonomous systems when blocked. Cloudflare says this second stream was observed across tens of thousands of domains and in some cases did not fetch robots.txt at all. When the declared crawler was refused, the undeclared one showed up.

The response was administrative. Perplexity lost its status as a verified bot, Cloudflare added fingerprinting heuristics to its managed rules so that customers who block AI crawlers now catch the stealth traffic too, and the company noted that its bot management system had already been scoring the undeclared requests as automated. For contrast the post ran the same test against OpenAI's ChatGPT-User agent, which fetched robots.txt, stopped when disallowed, and did not come back under another name.

What robots.txt was ever for

The Robots Exclusion Protocol dates from 1994 and was designed for a web of a few thousand sites and a handful of crawlers whose operators knew each other. It has no authentication, no enforcement, and no concept of purpose. A crawler declares a name, reads a text file, and decides whether to obey. For thirty years that mostly worked because the crawlers that mattered were search engines, and a search engine that ignored robots.txt would lose the goodwill it needed to be linked to.

The incentive that held the protocol together was reciprocity. Sites let crawlers in because crawlers sent visitors back. An AI answer engine breaks that loop. It fetches the page, answers the question, and the visitor never arrives. Cloudflare's own July figures put the referral gap at hundreds to tens of thousands of times worse than classic search for AI companies. Once the crawler gets nothing from being polite except the moral satisfaction, the text file is asking for restraint from an actor that has every reason not to show it.

Perplexity's likely defence, and we have seen versions of it made in public before this week, is that fetching a page on behalf of a user who asked a question is a user agent's action, not a crawler's, and so robots.txt does not apply. Whatever one thinks of that argument, it illustrates the problem. The protocol has no way to encode the distinction, so both sides can claim to be following it.

Why this matters for open data

For people who build datasets in the open, this incident is bad news on two fronts. The first is that the tools sites reach for when robots.txt fails are indiscriminate. Fingerprinting and managed rules that block stealth crawlers do not distinguish a research crawler from a commercial one, and a small academic crawler that does everything right will still look, to a heuristic classifier, like an unverified bot. The bar for access is moving from a text file anyone can read to a verification programme run by a private company.

The second is the effect on the corpora we can already see. Cloudflare says more than 2.5 million sites have used the tools it launched on July 1 to restrict AI training access. A crawl taken today reaches a different web from one taken in 2023, and it reaches it through a gate that records who is asking. Open datasets that document their crawl as a date and a seed list will need to add what verification the crawler carried and what fraction of the seeds refused it.

There is one small benefit. The test Cloudflare ran is reproducible by anyone with a domain and a firewall. Register a fresh site, disallow everything, ask the answer engines about it, and see who knows. That is a better compliance check than any user agent string, and we would like to see it run regularly and the results published.

What replaces the text file

The candidate replacement is already visible in the pay per crawl design from July. Crawlers sign requests with a key pair, register with the edge network, and either get a price or get refused. That system has authentication and enforcement, which robots.txt never had. It also has a gatekeeper, which robots.txt never had either. The web is trading a protocol nobody could enforce for one that a single vendor can, and the vendor gets to decide what counts as a verified bot.

We do not think there is a way back to the old arrangement. The reciprocity that held it together is gone. What we would want from whatever comes next is that the verification standard be open, so that a university crawler can register on the same terms as a company, and that the refusals be logged somewhere researchers can read them. The alternative is a web where the only people who can measure what was crawled are the people selling the gate.

Sources

  1. Cloudflare, Perplexity is using stealth, undeclared crawlers to evade website no-crawl directives
  2. Cloudflare, Content Independence Day