Two rulings in one week

On June 23 Judge William Alsup ruled in Bartz v. Anthropic that training a language model on lawfully acquired books is fair use, while holding that the seven million or so pirated copies Anthropic had assembled into a central library were a separate matter that would go to trial on damages. On June 25 Judge Vince Chhabria, in the same Northern District of California, granted Meta summary judgment on the claim that training Llama on thirteen authors' books infringed their copyright. Two wins for defendants in three days, and the second one reads like a brief for the plaintiffs.

The Meta order is worth reading in full because its structure is unusual. Four pages of general reasoning about why training on copyrighted books will often be unlawful, thirty pages of analysis, and a conclusion that the plaintiffs lost because they made the wrong argument. Chhabria says this directly on page five. The ruling, he writes, stands only for the proposition that these plaintiffs made the wrong arguments and failed to develop a record in support of the right one, and it should be clear that it does not establish Meta's use as lawful.

The argument the plaintiffs made

The authors, who include Sarah Silverman, Richard Kadrey, Matthew Klam and Rachel Louise Snyder, offered two theories of market harm. First, that Llama can reproduce snippets of their books. Second, that Meta's copying damaged the market for licensing books as training data. Chhabria calls both clear losers. Llama cannot generate enough of any plaintiff's book to matter, and there is no entitlement to a licensing market for a use that is otherwise fair. Meta, for its part, put in expert evidence that Llama 3's release had no discernible effect on the plaintiffs' sales.

The theory the court thought could win, that a model trained on a genre of books can flood the market with competing works and dilute the market for the originals, appears nowhere in the complaint and nowhere in the plaintiffs' own summary judgment motion. It surfaces as a fleeting reference in one expert report mentioning that AI-generated books are appearing on Amazon. The court lists the questions that reference leaves open, whether Llama can produce such books, whether they compete with a memoir or a story collection, and what the effect on sales would be compared to a world where developers cannot copy at all. On a record with none of that, speculation is not enough to reach a jury, and the fourth factor goes to Meta.

The argument the court wanted to hear

What makes the opinion a warning is how much of it is spent on the argument nobody made. The opening pages walk through biographies, magazine articles, romance and spy novels, and explain why a product that generates endless amounts of similar work would diminish the market for the originals and the incentive to write them. The court accepts that training is transformative as a factual matter and says that does not settle anything, because under fair use harm to the market is more important than the purpose of the copying. The upshot, in the court's words, is that in many circumstances it will be illegal to copy protected works to train generative models without permission, and that companies will generally need to pay for the right.

The order also takes a direct swing at Alsup's reasoning from two days earlier. Alsup had compared training to teaching schoolchildren to write, which might result in an explosion of competing works, and found that this was not the kind of displacement copyright cares about. Chhabria calls the analogy inapt. Using books to teach children is not remotely like using books to build a product that lets a single person generate countless competing works with a fraction of the effort. He is equally dismissive of the argument that adverse rulings would halt the technology, noting that Meta itself had tried to license books before finding it too logistically difficult.

The conclusion goes further than dicta usually does. In cases involving uses like Meta's, the court writes, plaintiffs will often win, at least where the record on market effects is better developed. It then sketches the exceptions, nonprofit uses like national security or medical research, and plaintiffs whose works face no meaningful AI competition. That is a roadmap, and it was written by a judge who had just ruled against the people who could have used it.

What each ruling left open

Neither case is over. In Bartz the trial on the pirated library remains, and the exposure there is statutory damages on millions of works, which is the number every developer with a shadow library download in its history now has to think about. In Kadrey the court granted Meta judgment on the DMCA claim separately and set a July 11 conference on the plaintiffs' remaining claim that Meta distributed their works while torrenting them.

The two orders also disagree on what the shadow library download means. Alsup treated the pirated copies as an independent wrong. Chhabria treated the download as part of the training use, to be judged by its ultimate purpose, and said the loss of isolated sales to developers is not the kind of harm that tips the scale. He did leave open whether Meta's use benefited the shadow libraries themselves, and noted that the plaintiffs did not argue it.

What this means for open training data

Both rulings bind nobody outside their courtrooms, and Kadrey is explicitly not a class action, so its effect is limited to thirteen authors. But the reasoning will be quoted in every complaint filed from here on, and the next complaint will plead market dilution on page one with an economist attached. The EFF has already argued that Chhabria's dilution theory mistakes competition for infringement, and we expect that fight to reach an appellate court within two years.

For people who build open corpora, the practical reading is narrow and clear. Acquisition matters as much as use. Alsup drew the line at piracy. Chhabria would draw it at market effect but found the piracy relevant to the fourth factor if anyone had argued it. A corpus assembled from purchased or licensed sources with a manifest is defensible under both opinions. One assembled from a shadow library mirror is exposed under one and an open question under the other. What we would like to see next is an actual measurement of the dilution theory, meaning a study of whether AI-generated books in a genre have moved sales of human-written ones. Chhabria asked for the evidence. Someone in our field should try to produce it before a court does.

Sources

  1. Kadrey v. Meta Platforms, order on cross-motions for partial summary judgment, N.D. Cal. No. 23-cv-03417-VC, June 25, 2025 (CourtListener)
  2. Wikipedia, Bartz v. Anthropic
  3. Wikipedia, Artificial intelligence and copyright
  4. EFF, Two courts rule on generative AI and fair use