Content Independence Day: Cloudflare flips the crawl default
On July 1 Cloudflare started blocking AI crawlers by default for new customers and opened a private beta that answers crawlers with HTTP 402. Notes on what a billable web means for anyone who builds open datasets from it.
What changed on July 1
Cloudflare announced on July 1 that new customers will have AI crawlers blocked by default, and that a crawler wanting access to a site will have to be allowed explicitly or pay for it. The company framed the date as Content Independence Day and said it had lined up publishers and AI companies behind the change. The mechanism for the paying part is a new product called pay per crawl, currently in private beta.
The argument Cloudflare made for the change is about referral traffic. By its figures, getting traffic from Google today is ten times harder than it was from the original Google search, getting traffic from OpenAI is 750 times harder, and getting traffic from Anthropic is 30,000 times harder. The post also cites a figure that 75 percent of mobile queries are now answered inside Google without the user leaving the page. Whatever one thinks of the exact ratios, the direction is the same one every publisher has been describing for two years. Crawlers take the content, and the visit that used to come back does not.
How pay per crawl works
The design is small and worth understanding, because it will shape what the next generation of web corpora looks like. A crawler that wants a page either declares a willingness to pay up front or gets refused with a price. In the reactive flow, the crawler requests a page and receives HTTP 402 Payment Required with a crawler-price header. It can retry with a crawler-exact-price header matching that price and get a 200 response. In the proactive flow, the crawler sends crawler-max-price on the first request, and a successful charge comes back confirmed in a crawler-charged header on the 200 response.
Publishers set one flat per-request price for the whole domain, and for each crawler choose allow, charge, or block. Cloudflare acts as merchant of record and settles the money. Crawlers have to identify themselves cryptographically through Web Bot Auth, which means an Ed25519 key pair and HTTP message signatures carried in signature-agent, signature-input, and signature headers, and they have to register with Cloudflare before they can receive a payment request at all. The 402 status code has sat in the HTTP specification unused for most of its life. This is the first time we have seen it wired to a real settlement system at scale.
Why this matters for open datasets
Open pretraining corpora, from Common Crawl derivatives to the filtered web sets research groups publish, have always rested on an assumption that the public web is crawlable by anyone who respects robots.txt. That assumption was already fraying as sites added no-crawl directives. What changed this week is the default. A site that does nothing now blocks, and the path to access runs through a registered identity and a price.
Two consequences follow. First, the composition of any corpus built after July 2025 will differ systematically from one built before, because the sites that stay open will be the ones that chose to. Anyone comparing models trained on old and new crawls will need to account for that shift, and we do not think we yet know which kinds of content drop out. Second, a paid crawl is a crawl with a record. If a lab pays for pages, there is a ledger of what it fetched, which is a form of provenance the field has never had. Whether that ledger is ever made public is a different question.
The scheme also creates a cost asymmetry that cuts against the people who release data openly. A company that trains a proprietary model can recover crawl fees from its product. A group that assembles a corpus to give away cannot, and it will face the same 402 responses. Nothing in the announcement distinguishes research crawling from commercial crawling, and the registration requirement means anonymous academic crawlers will be treated as unverified bots.
What we would watch
The obvious question is whether crawlers comply. A crawler that ignores the signal by presenting itself as a browser does not receive a 402, it receives whatever the bot detection decides to serve. The whole system rests on Cloudflare's ability to tell declared crawlers from undeclared ones, and on its willingness to punish the undeclared. We would expect the first public dispute over a stealth crawler within months.
The second question is price discovery. A flat per-request price for a whole domain is crude, and Cloudflare's own post says it expects dynamic pricing and finer licensing to follow. If a market forms, the price of a page will tell us something no benchmark does, which is what training data is actually worth to the people buying it. We would like someone to collect those prices as they become visible and publish the series.
For our own work, the practical change is that we should stop treating a crawl date as a minor detail in a dataset card. It is now a policy regime. Corpora built before and after July 1, 2025 were gathered under different rules, and the difference should be recorded alongside the licence.
Sources
From the foundation