What happened

On February 11 Judge Stephanos Bibas, a Third Circuit judge sitting by designation in the District of Delaware, granted summary judgment to Thomson Reuters on fair use in its copyright case against Ross Intelligence. The opinion opens by admitting error and revising his own 2023 opinion, which had sent most of the case to a jury. It is one of the first US decisions on whether using copyrighted material to train an AI system is fair use, and it went against the AI company.

Westlaw, owned by Thomson Reuters, publishes headnotes, short editorial summaries of the key points of law in a judicial opinion. Ross was building a legal research search engine to compete with Westlaw. It asked to license Westlaw content and was refused because it was a competitor. So Ross bought roughly 25,000 Bulk Memos from a company called LegalEase, compilations of legal questions with good and bad answers, and lawyers writing those memos had been told to build the questions using Westlaw headnotes. Ross trained on the memos, Thomson Reuters sued in 2020, and Ross went out of business in 2021 while the case continued.

The copyright side

Before reaching fair use the court had to decide whether headnotes are copyrightable at all, and here Bibas changed his mind. In 2023 he thought a jury would need to decide, since a headnote that closely tracks the opinion it summarises might not be original. Now he holds that choosing which point of an opinion to distil into a headnote is itself the creative spark, like a sculptor choosing what to carve away, so even a headnote close to the opinion text clears the low originality bar from Feist.

He then went through the accused headnotes one by one. Thomson Reuters had accused Ross of infringing 21,787 headnotes, but the batch before the court on this motion was 2,830. Bibas compared each Bulk Memo question against the headnote and the underlying opinion, and granted summary judgment on 2,243 of them, where the question looked more like the headnote than like the opinion and no reasonable jury could find otherwise. The rest go to trial.

The four factors

Factor one, the purpose and character of the use, went to Thomson Reuters. Ross conceded its use was commercial. The question was whether it was transformative, and the court applied the Supreme Court's Warhol decision, under which a secondary use that shares the same purpose as the original and is commercial will usually fail the factor. Ross was using headnotes as training data to build a legal research tool that competes with Westlaw. Same purpose, same market. Not transformative.

Ross's best argument was intermediate copying. The headnotes never appear in the product. They were consumed at a training step and what the user sees is a list of judicial opinions. Ross pointed to the Sega and Sony v. Connectix line of cases, where copying software to reverse engineer it was fair use. Bibas rejected the analogy on two grounds. Those were computer programming cases about getting access to unprotected functional elements, and the copying in them was necessary for a competitor to build something new. Neither applies to headnotes. In 2023 he had relied on those same cases, and he says now that was a mistake.

Factors two and three went to Ross. The nature of the work factor favoured Ross because headnotes are not highly creative, though the court notes that factor rarely decides anything. The amount used factor also favoured Ross because the headnotes do not reach the public through Ross's product. Factor four, market effect, went to Thomson Reuters and Bibas calls it the most important element. The original market is legal research platforms, and Ross set out to be a market substitute. There is also a potential derivative market in data to train legal AIs, and the court holds it does not matter whether Thomson Reuters has actually used its data that way, because the effect on the potential market is enough and Ross carried the burden of showing that market did not exist. Weighing factors one and four most heavily, fair use fails.

The sentence that limits everything

In the middle of the factor one analysis, the opinion says that it is undisputed that Ross's AI is not generative AI, meaning AI that writes new content itself. When a user asks Ross a legal question, it returns existing judicial opinions. Bibas adds, for readers, that because the AI field is changing rapidly he wants to note that only non-generative AI is before him. The Wikipedia summary puts it plainly. Ross's output was composed of pieces of Westlaw's material, and how the case applies to generative systems is not clear.

That caveat is doing a great deal of work. The transformative use argument for a language model is that it learns statistical structure and produces new text rather than retrieving the inputs. Ross could not make that argument because it was a retrieval system trained on near copies of the inputs. The market substitution argument was easy because Ross was, by its own admission, trying to replace Westlaw in the same market. A general purpose model trained on a broad web corpus is not a direct substitute for any one plaintiff's product in the same way. So the decision most helpful to rightsholders is the one whose facts are least like the cases that matter most.

What the ruling does shape

Two pieces of reasoning travel regardless of the generative distinction. The first is the derivative market holding. Bibas found a potential market for AI training data and said the plaintiff does not have to be selling into it already for harm to that market to count. Every rightsholder now has a template for arguing that the possibility of licensing deals establishes a market that unlicensed training damages.

The second is the rejection of intermediate copying as a general defence. Much of the informal case for training being fair use rests on the idea that copying inside a pipeline that never reaches a user is different in kind. This opinion says that the Sega line of cases is narrow, about code and about necessity, and that copying at an intermediate step is judged by the broad purpose of the use like any other copying. That applies to a language model just as well as to a search engine.

What we would watch for is the interlocutory appeal. The opinion is a partial summary judgment in a case that still has to go to trial on the remaining headnotes, and the fair use question is a pure question of law here because the facts were undisputed. If the Third Circuit takes it, the first appellate word on AI training will come from a case with a defunct defendant and a non-generative system, and everyone litigating the generative cases will have to argue around it.

Sources

  1. Memorandum Opinion, Thomson Reuters v. Ross Intelligence, D. Del., February 11, 2025 (CourtListener RECAP)
  2. Wikipedia: Artificial intelligence and copyright