What the complaint asks for

The New York Times filed suit against Microsoft and OpenAI in the Southern District of New York on December 27, docket 1:23-cv-11195, assigned to Judge Sidney Stein. The complaint alleges copyright infringement and trademark dilution. It says the defendants trained models on millions of Times articles without authorization, that ChatGPT and Copilot reproduce Times reporting, and that this competes with the paper and costs it subscribers and revenue. The Times had been in licensing talks with OpenAI, which had signed deals with Axel Springer and the Associated Press but not with the Times.

The remedies are the part that matters for anyone in our field. Beyond damages, which the complaint describes as billions without naming a figure, the Times asks for destruction of the datasets and models that incorporate its work. That is a request for the court to order a model deleted. We have not seen that asked for before in a case of this size, and it means the outcome will not just be a cheque.

The other thing the complaint does differently is show its evidence. Where earlier suits argued from the fact of training, the Times exhibits show articles reproduced almost verbatim in model output. Dinusha Mendis, writing at The Conversation, contrasts this with Getty v. Stability, where the outputs were derivative rather than regurgitated. Regurgitation is a stronger fact for a plaintiff and a harder one for a fair-use defence to absorb.

The fair-use argument, as it stands

OpenAI's position is that training is transformative under Section 107 and that the suit is without merit. Its January 8 response also says training leading models would be impossible without copyrighted material, which is an argument about necessity rather than about law, and we doubt a court will find it decisive either way.

The Times' counter is about the fourth factor, market harm. The complaint says the chatbots bypassed paywalls to produce article summaries, which if true is a direct substitute for the product the Times sells. Mendis notes that this kind of commercial impact is exactly what undermines a fair-use claim. We are not a lawyer and will not predict the ruling. But the shape of the argument is now visible, and it turns on what the model emits, not only on what it ingested.

That shift matters technically. It means memorization is a legal liability, not just a modelling defect. A lab that can show its model does not reproduce long passages of training text has a better defence than one that cannot. Measuring that is an open research problem, and until this week it was mostly an academic one.

What this does to open corpora

Here is the problem for open science. The labs with the most exposure in this suit train on undisclosed data. That opacity is now a defensive asset, because a plaintiff cannot exhibit what it cannot see. A group that publishes its training corpus in full, as we would like to, is handing every rights holder a searchable index of their own work. The incentive runs directly against disclosure.

We do not think the answer is to stop publishing corpora. We think the answer is to publish corpora that can withstand the search. That means a manifest with a source and a licence per document, a stated policy on paywalled and robots-excluded content, and a memorization audit on the trained model that reports how much verbatim training text it can be induced to produce. None of that is free, and all of it is cheaper than being the test case.

We have a small taste of this ourselves. Our web corpus excludes robots-excluded domains and anything behind a paywall, which removes most major newspapers by construction. Nobody had to sue us for that. It was a choice made because we intended to publish the manifest and did not want to publish a list of sites we had no right to be on. The Times' complaint is a reminder that the choice has a cost in coverage of current events, and that the cost was there whether or not we measured it.

It also means being honest that a fully documented corpus will be smaller and older than an undocumented one. If the court finds for the Times, that gap becomes the price of a legal model and everyone pays it. If the court finds for OpenAI, the gap remains the price of a reproducible one, and only people who care about reproducibility pay it. Either way it is worth knowing how large the gap actually is, and nobody has measured it.

What we would want someone to try

The experiment we want is straightforward and expensive. Train two models of the same size and token budget, one on a corpus with news content included and one with it excluded, and report the difference on the benchmarks people cite plus a memorization audit on both. That would tell us how much of the capability that is now in dispute came from the disputed text.

Second, the field needs a standard for the memorization audit itself. Right now every lab that reports extraction rates uses its own prompts and its own thresholds. A shared protocol, published openly and run by a third party, would give both sides of this case something better than screenshots to argue from.

The suit will take years. The habit of quiet training data is over now, and whatever we build in the meantime should assume that someone will eventually ask to see the list.

Sources

  1. Wikipedia: The New York Times v. Microsoft and OpenAI
  2. The Conversation: How a New York Times copyright lawsuit against OpenAI could transform how AI and copyright work