A report that was not supposed to look like this

The third part of the Copyright Office's report on artificial intelligence went up on May 9 with the words pre-publication version stamped on the cover. Part 1, on digital replicas, had come out in July 2024 and Part 2, on copyrightability, in January 2025, both as finished documents. Part 3, on generative AI training, arrived with a note on the inside cover saying the Office was releasing it in response to congressional inquiries and expressions of interest from stakeholders, and that a final version would follow without any substantive changes expected in the analysis or conclusions.

The next day, May 10, the Register of Copyrights, Shira Perlmutter, was dismissed by the administration. The Librarian of Congress who had appointed her, Carla Hayden, had been removed earlier the same week. Perlmutter has since filed suit challenging the dismissal, and the question of who runs the Office is now in the courts. We want to separate what the report actually concludes from the noise around it, because the two are being conflated in most of the coverage we have read.

What the report concludes

The document is 108 pages and draws on more than 10,000 comments submitted in response to a notice of inquiry the Office issued in August 2023. Its structure is conventional. It first asks whether the stages of building a generative model involve acts that implicate exclusive rights, and concludes that several do, including data collection and training. It then asks whether those acts can be excused as fair use, and it declines to give a single answer.

The core position is on page 74, in a section titled Weighing the Factors. The Office expects some uses of copyrighted works in training to qualify as fair use and some not. At one end, uses for noncommercial research or analysis that do not enable portions of the works to be reproduced in the outputs are likely to be fair. At the other, copying expressive works from pirate sources in order to generate unrestricted content that competes in the marketplace, when licensing is reasonably available, is unlikely to qualify. Many uses, the report says, will fall somewhere in between. The conclusion restates this in slightly stronger terms, saying that commercial use of vast troves of copyrighted works to produce expressive content that competes with them in existing markets, especially through illegal access, goes beyond established fair use boundaries.

Two analytic moves are worth flagging. On the first factor, the report accepts that training is often transformative, but insists, following Warhol, that transformativeness is a matter of degree and depends on what the model is deployed to do. On the fourth factor, it introduces market dilution as a cognisable harm. The argument is that the speed and scale at which systems generate content poses a serious risk of diluting markets for works of the same kind as the training data, so that if thousands of AI-generated romance novels appear, fewer human-written ones sell. The Office acknowledges this is uncharted territory and says the fourth factor should not be read so narrowly as to exclude it.

What it recommends

On licensing, the report is more cautious than either side wanted. It rejects a compulsory licence, saying such regimes fix rates and practices in stone and can take years to negotiate. It says an opt-out approach is inconsistent with the basic principle that consent is required for uses within the scope of statutory rights. It notes that voluntary licensing deals are emerging in music, images and news, and it recommends allowing that market to continue to develop without government intervention, with extended collective licensing to be considered only where a market failure is shown for specific types of works in specific contexts.

The report also addresses the argument, made by a16z and others, that treating training as infringement would entrench the largest companies because only they can afford to license. The Office says that concern should not be discounted but does not alter the fair use analysis, and that competition problems belong to antitrust law and the agencies that enforce it. That is a tidy separation, and whether it holds up in practice is a different question.

Why the release matters as much as the text

The report itself is an analytical framework with no binding force. It says so on page 2, that it provides an analytical framework without opining on specific cases, and that it draws on the Office's experience advising Congress and the courts. Courts are not bound by it. What it does is set out, in one place and with a full record behind it, a reading of fair use that neither treats training as automatically fair nor treats it as automatically infringing. Litigants on both sides of the dozens of pending cases will cite the parts they like.

The pre-publication label changes how that citation will go. A final report would carry the institutional weight of the Office. A draft released in a hurry, followed the next day by the removal of the person who signed it, carries an asterisk. Anyone arguing against the market dilution theory now has a procedural line to add, that the document was never finalised and its author was dismissed. Anyone arguing for it will point out that the Office promised no substantive changes. The text has not changed. Its authority has, and that shift happened outside the document.

What we would watch

Three things. Whether a final version of Part 3 ever appears, and under whose signature. Whether any court adopts the market dilution framing of the fourth factor, which is the most novel idea in the report and the one most exposed to the argument that it is policy rather than law. And whether the research carve-out at the fair end of the spectrum, noncommercial analysis with no reproduction in outputs, gets read narrowly or broadly, because that is the line most of the academic training runs we know about are standing on.

Sources

  1. U.S. Copyright Office, Copyright and Artificial Intelligence, Part 3: Generative AI Training (pre-publication version)
  2. U.S. Copyright Office, Copyright and Artificial Intelligence report page
  3. Wikipedia, Shira Perlmutter