What was filed

On Tuesday the Authors Guild and 17 novelists filed a class action against OpenAI in the Southern District of New York. The named plaintiffs include John Grisham, George R.R. Martin, Jodi Picoult, David Baldacci, Jonathan Franzen, George Saunders and Scott Turow. The proposed class is fiction writers whose works were used to train GPT models. The complaint runs 47 pages and asks the court to stop the unauthorised use and to compensate authors.

The Guild's own announcement puts the provenance claim in its second paragraph. The complaint, it says, charges that the books were downloaded from pirate ebook repositories and then incorporated into GPT-3.5 and GPT-4. Co-counsel Rachel Geman is quoted saying that without the plaintiffs' works the defendants would have a vastly different commercial product.

This is the third author suit against OpenAI this summer. Tremblay v. OpenAI was filed in the Northern District of California on June 28, and a companion case from Sarah Silverman and others followed. The new filing recruits far bigger names and moves the venue to New York. The more interesting difference is in which allegation carries the weight.

Books1 and Books2, and how the estimate is built

Both the June complaint and this one start from the same eight-page passage in OpenAI's 2020 GPT-3 paper, which lists the training mix as Common Crawl, WebText2, Wikipedia and two internet-based books corpora it calls Books1 and Books2. OpenAI never said what was in them. Everything the plaintiffs allege about their contents is inference from size.

The Tremblay complaint does the arithmetic explicitly. Taking BookCorpus at about 7,000 titles, and using the token counts in the GPT-3 paper, it estimates Books1 at about nine times that size, roughly 63,000 titles, and Books2 at about 42 times, roughly 294,000 titles. It then guesses Books1 is Project Gutenberg, which had just over 60,000 public domain titles in 2020, and that Books2 was copied from shadow libraries such as Library Genesis, Z-Library, Sci-Hub and Bibliotik, whose collections circulate in bulk over torrents.

The Authors Guild complaint takes the same route with different numbers. Paragraph 102 says the disclosed size of Books2, 55 billion tokens, implies more than 100,000 books. Paragraph 103 notes that Books3, the dataset an independent researcher compiled from Bibliotik, holds nearly 200,000 books and has been used by other developers. Paragraph 104 argues that similar sizes plus the small number of sites that allow bulk ebook downloads strongly indicate Books2 came from one of those repositories.

Why this allegation carries the case

Read the paragraphs in order and the structure is clear. Paragraph 96 says OpenAI refuses to discuss the source of Books2. Paragraph 105 says it has not discussed the data for GPT-3.5 or GPT-4 at all. Paragraph 107 argues that models this much larger must have used a correspondingly larger book corpus, and 108 states flatly that there is no other way OpenAI could have obtained the volume of books required. Paragraph 110 then alleges the copying was wilful.

The reason this matters more than the output claims is that the output claims got weaker over the summer. The complaint itself records, in paragraphs 88 to 91, that ChatGPT used to return verbatim excerpts of copyrighted books and now declines, offering summaries instead. Summaries that contain details unavailable in reviews are offered as evidence the full text was ingested, and the complaint lists generated outlines with invented titles for sequels in the plaintiffs' worlds. But summaries are a thinner infringement theory than verbatim reproduction, and every lab has learned to suppress verbatim output.

Provenance does not have that problem. If discovery shows that Books2 was a LibGen or Bibliotik dump, the question of whether training is fair use gets tangled with the question of whether the copy used to train was itself lawfully obtained. The complaint points out that LibGen is already known to this court from the Elsevier v. Sci-Hub case. That is a very different posture from arguing about what a chatbot says when asked to write in Grisham's voice.

What the complaint does not have

It does not have a document. Every provenance claim is on information and belief, resting on token counts, on the size of Books3, and on the observation that there are few places to get 100,000 ebooks. The suit is designed to survive a motion to dismiss and then get the answer in discovery. Whether a court lets plaintiffs go looking inside a training pipeline on the strength of an inference from corpus size is the procedural question that will decide a lot.

It also does not distinguish between GPT-3, where the Books1 and Books2 names come from, and GPT-3.5 and GPT-4, where nothing is disclosed. The complaint bridges the gap by arguing that more parameters imply more books, which is plausible and unproven. The parameter figures it cites for GPT-4 are themselves press estimates rather than anything OpenAI has said.

What this means for anyone who trains models

The lesson we take from the filing is that the dataset itself has become the legal object. The Guild's FAQ says it uses the word train only because it has become shorthand, and that the works are used to build the AI and remain part of its fabric. Whatever one thinks of that framing, the litigation strategy follows from it. Establish where the copies came from and the model inherits the problem.

For a research group that publishes its data, the implication is direct. A provenance record for every corpus, with the source, the licence, and the acquisition date, is no longer good hygiene. It is the thing you will be asked for first. The labs that cannot produce one are about to find out what it costs to answer the question in court instead of in a README.

Sources

  1. The Authors Guild, John Grisham, Jodi Picoult, David Baldacci, George R.R. Martin, and 13 Other Authors File Class-Action Suit Against OpenAI (Authors Guild, September 2023)
  2. Authors Guild v. OpenAI, complaint, No. 1:23-cv-08292 (S.D.N.Y., filed September 19, 2023)
  3. Tremblay v. OpenAI, complaint (N.D. Cal., filed June 28, 2023)
  4. Authors Guild, AI FAQ