The deal

The press release went out on May 6. OpenAI gets access to OverflowAPI, the product Stack Overflow built to sell its corpus of vetted technical questions and answers to model developers. In return, ChatGPT will surface Stack Overflow content with attribution back to the community, Stack Overflow will use OpenAI models inside its own OverflowAI product, and the company says it will reinvest the proceeds in community features. First integrations were promised for the first half of this year.

The quotes are what you would expect. Brad Lightcap, OpenAI's COO, said that learning from as many languages, cultures, subjects and industries as possible ensures the models can serve everyone. Prashanth Chandrasekar, Stack Overflow's CEO, pointed to more than 59 million questions and answers. Neither statement is false. Neither addresses the question that the people who wrote those 59 million posts started asking within hours.

The protest and the response

The protest took the only form available to a contributor with no vote on the deal. People went to their highest-voted answers and either deleted them or edited them into a statement objecting to the partnership. One user, identified as Ben, described on Mastodon what happened next. He changed his top answers to protest messages, moderators reverted them within the hour, and his account was suspended for a week.

Tom's Hardware reported the pattern across the site by May 8. Deletion of well-received answers, particularly accepted ones, is blocked by the platform's rules. Edits that replaced content with protest text were rolled back. Accounts that persisted were suspended, with seven days the documented length. Stack Overflow's position, as reported, is that contributions are licensed to the site irrevocably under Creative Commons and that removing an accepted answer damages the knowledge base other people rely on.

What the licence actually says

The licensing page is short and settles the legal question without settling the moral one. Contributions are licensed under Creative Commons Attribution-ShareAlike, at version 2.5 for anything before April 2011, 3.0 through May 1, 2018, and 4.0 since. A CC BY-SA licence is perpetual and cannot be withdrawn once granted. Anyone, including OpenAI, has been free to use every Stack Overflow post for years, provided they attribute and share alike.

That is the awkward part for the protesters. The licence they agreed to already permitted the use they are objecting to. What the OpenAI deal adds is a paid, structured feed and a commercial relationship, and what it removes is the pretence that share-alike means anything once the content goes into a model. A model trained on CC BY-SA text does not itself carry a CC BY-SA licence, and nobody has forced the issue in court. The attribution ChatGPT promises is a product feature, offered voluntarily, and the deal is structured so that it does not have to be a legal obligation.

Who owns a commons

Stack Overflow was built on a bargain that was never written down. You answer questions for free, and in exchange the answers stay public, free, and credited to you. The company hosts them, sells advertising and job listings around them, and keeps the lights on. For fifteen years that bargain held because the value of the corpus was inseparable from the site that displayed it. A model changes that. The corpus has value on its own, as training data, and the company that holds the servers can sell that value without the people who created it.

The same year the site cut its own legs off. The reversal of its ban on AI-generated answers, and the moderator strike that followed in 2023, were the first sign that the company and the community had different ideas about what the site was for. This deal is the second. The contributors thought they were building a commons. The company's actions say it was building an asset.

We do not think the protesters have a legal case and we do not think that matters much. What they have demonstrated is that the supply of high-quality answers depends on people who now know what happens to their work. If a significant fraction of them stop, the OverflowAPI feed gets thinner every year, and there is no amount of licensing text that can fix that.

What this means for anyone publishing openly

For a foundation that publishes everything in the open, this is not an abstract question. Every dataset and paper we release goes into someone's training run, and the licences we choose will not stop that. What we can do is be honest about it in advance, choose licences that reflect what we actually want, which for us is reuse, and avoid the surprise that Stack Overflow's contributors just had.

The other lesson is about platforms. Contributions to a site owned by a company are governed by that company's terms, and the terms can change or be interpreted in new ways when a buyer appears. The protest failed because the platform controls the delete button. Work that lives in a repository you control, under a licence you chose, with a copy somewhere the platform cannot reach, cannot be sold out from under you in quite the same way.

Sources

  1. Stack Overflow press release, Stack Overflow and OpenAI Partner to Strengthen the World's Most Popular Large Language Models (May 6, 2024)
  2. Tom's Hardware, Stack Overflow bans users en masse for rebelling against OpenAI partnership (May 8, 2024)
  3. Stack Overflow help centre, What is the licensing on user contributions?