What was audited

Shayne Longpre and 48 co-authors from the Data Provenance Initiative posted this audit on July 20. They took three of the pretraining corpora that open models are actually built on, C4 with 15.9 million domains from an April 2019 crawl, RefinedWeb with 33.2 million domains, and Dolma with 45.2 million, and studied 14,000 domains from them. Two samples matter. HeadAll is the top 2,000 domains by token count in each corpus, 3,950 unique domains in total. Random2k is a random draw from the intersection.

For each domain they read the robots.txt over time and the terms of service, and coded what each one permitted, for which crawlers. The distinction between the two mechanisms turns out to be one of the paper's main results, so it is worth holding onto. Robots.txt is machine-readable and rigid. Terms of service are expressive and nobody's crawler reads them.

The numbers for one year

Between April 2023 and April 2024, the fraction of C4 tokens fully restricted by robots.txt went from about one percent to between five and seven percent. In the HeadAll sample, the high-quality, actively maintained sites that carry the most tokens, the restricted fraction went from under three percent to between 20 and 33 percent. The abstract's headline figure is that over 28 percent of the most critical sources in C4 became fully restricted in a single year. Terms of service restrictions on crawling now cover 45 to 55 percent of all tokens.

News drove the change. By April 2024 about 45 percent of tokens from news domains were fully restricted by robots.txt, against three percent a year earlier. The authors ran a SARIMA forecast and expect another two to four points of robots.txt restriction on the full corpus by April 2025, and six to ten more points of terms of service restriction. Whether one trusts a time series forecast on 13 months of data, the direction is not in doubt.

Who the blocks name

Here is the finding that changed how we think about it. Among domains that restrict any AI crawler at all, 91.5 percent restrict OpenAI's agent, 83.4 percent restrict Common Crawl, 83.4 percent restrict Anthropic, 72.0 percent restrict Google-Extended, and around 52 percent restrict Cohere and Meta. The Internet Archive is restricted on 32.3 percent of those domains, and Google Search on 17.1 percent. The authors' reading is that administrators either do not know the smaller crawlers exist or cannot keep a list current, so they name the two or three they have heard of, or they use a wildcard.

The wildcard is the problem for us. A site that writes a blanket disallow to keep out one commercial lab also keeps out Common Crawl, which the paper reports is cited in more than 10,000 research articles, and the Internet Archive. The authors say directly that the rising expression of non-consent will affect non-profits, archives and academic researchers. The commons is closing fastest for the people who never had a way to pay for access in the first place, because the sites that chose to negotiate with a lab were negotiating with a company, and a wildcard does not distinguish a lab from a university.

Two signals that disagree

The paper finds that 35.1 percent of domains have terms of service that forbid crawling but no robots.txt restriction, and 20.3 percent have a restrictive robots.txt but no terms of service. The two mechanisms frequently disagree about what they express and about what they are able to express. The authors trace this to the Robots Exclusion Protocol having been designed in 1995 for a different web. It can say which agent may fetch which path. It cannot say that fetching for search is fine and fetching for training is not, which is the distinction most site owners now want to draw.

There is a related mismatch between what corpora contain and what models are used for. News is around 40 percent of the HeadAll tokens in C4 but under one percent of ChatGPT queries in the WildChat data the authors compare against, while creative writing and role-play make up more than 30 percent of real use and are thin in web corpora. The paper raises this in the context of fair use analysis, but we read it as a data quality point too. The sources that are closing are not the ones users are drawing on most.

What a researcher should do with this

The practical consequence for anyone building an open corpus this year is that crawls respecting robots.txt will drift toward organisation and e-commerce sites and away from news, forums and social media, and that the drift will be invisible unless someone measures it. A model trained on a 2025 crawl and a model trained on a 2023 crawl will differ in composition for reasons that have nothing to do with the web changing and everything to do with who got blocked.

What we would want next is a machine-readable way for a site to say yes to research and no to commercial training, and a study of whether sites would use it. The authors argue for better protocols. We agree, and we think the evidence here says the demand exists, since a third of restricting domains already carve out the Internet Archive when they can. Until such a protocol exists, every open corpus should ship with the date of its crawl and the fraction of its head domains that have since closed.

Sources

  1. Longpre et al., Consent in Crisis: The Rapid Decline of the AI Data Commons (arXiv 2407.14933)
  2. Full text of the paper (arXiv HTML)