Our content is free, our infrastructure is not: Wikimedia counts the crawlers
Wikimedia reported that bots generate 65 percent of its most expensive traffic and that multimedia bandwidth has risen 50 percent since January 2024. On the hidden cost open knowledge commons pay to feed training runs, and the access channels Wikimedia is proposing instead.
The numbers
On April 1 the Wikimedia Foundation published a post with two figures that anyone training on web data should read. Bandwidth for multimedia content has grown 50 percent since January 2024, and the growth is coming from scrapers rather than readers. And while bots account for about 35 percent of pageviews, they account for 65 percent of the traffic that reaches the core datacenters, which is the most expensive kind.
The mismatch between those two percentages is the whole story. Wikimedia serves popular pages from caches close to readers. Human attention is concentrated, so most human requests hit a cache and cost little. Crawlers read everything, including pages nobody has looked at in years, and those requests miss the cache and go all the way back to the origin. A bot reading the long tail costs the Foundation far more per request than a person reading the front page.
What a real spike looks like
The post gives a control case. When Jimmy Carter died in December 2024, his article received 2.8 million views in a day, and many people also streamed a 1.5-hour video of one of his presidential debates. Network traffic doubled and some connections filled for about an hour, slowing page loads for some users until the reliability team rerouted traffic. That is what a genuine surge in human interest does to the infrastructure, and it was handled within an hour.
Crawler load is different in kind. It is a rising floor rather than a spike, and it is spread across the pages least likely to be cached. The Foundation says the growth in baseline multimedia demand shows no sign of slowing. The same pressure is hitting the developer infrastructure, the code review system and bug tracker, and the post describes engineering time increasingly going to blocking bots rather than to anything else.
Who pays for training data
Wikipedia is in nearly every pretraining corpus we have ever seen described. It is small in tokens compared with the web crawl but disproportionately valued, and Commons images are a large share of the openly licensed multimedia available anywhere. All of that is free to use under its licence, and the licence has always been the point. What the licence does not cover is the bandwidth and the servers, which are paid for by donations.
So the arrangement as it stands is that a knowledge commons funded by small donors is absorbing the operating cost of data collection for companies whose training runs cost hundreds of millions. Wikimedia is not the only project in this position. Every open archive, every preprint server, every code host with a public API is seeing some version of the same curve. Wikimedia is one of the few with the monitoring to put a number on it and the standing to say it publicly.
The proposed alternative
The Foundation is not asking crawlers to stop. Its 2025 to 2026 annual plan includes an objective called Responsible Use of Infrastructure, WE5, whose aim is to establish sustainable channels for developers and reusers to get content without scraping the reader-facing site. The framing in the post is that the content is free and the infrastructure is not, and that the way to keep the first true is to set boundaries on the second.
What that should mean in practice is bulk access designed for machines, dumps and APIs served from infrastructure sized for the purpose, with the crawlers identified and, where they are commercial, contributing to the cost. Wikimedia already publishes full database dumps, which makes the scraping of the live site harder to defend. A crawler hitting every page of the long tail is choosing the expensive path when a cheap one exists, usually because the cheap path takes a little engineering to consume.
What we would like the labs to do
Three things, in order of difficulty. Identify your crawler honestly in the user agent and respect robots.txt, which costs nothing. Consume dumps and bulk APIs where they exist instead of crawling live sites, which costs a little engineering. And fund the infrastructure you depend on, in proportion to what you take from it, which costs money. The third is the one that would actually change the curve, and it is the one nobody has publicly committed to.
The number we want to see next year is the bot share of core datacenter traffic after WE5 has been running for a while. If it is still 65 percent, the polite route failed and the commons will start closing doors. If it has dropped, we will have a template that other open archives can copy.
Sources
From the foundation