What the court decided

Last week the German collecting society GEMA won its case against OpenAI in Munich. The court found that ChatGPT had returned copyrighted song lyrics almost identically when users asked for them. OpenAI's defence was that the texts were newly generated by the model rather than copied. The court did not accept that framing. Its reasoning turned on the fact that lyrics used in training came back nearly verbatim on request, which it took as evidence that the works had been stored in the model.

The remedies follow from that finding. OpenAI must cease saving the lyrics and cease returning them in its services, pay damages to GEMA, and provide usage data for the lyrics in question and the revenue generated from them. We have not been able to read the judgment itself, and we are working from secondary summaries, so we will keep to what those summaries agree on.

Why memorisation was the argument that worked

European copyright law already has an answer to the question of whether training on copyrighted text is allowed. The 2019 Directive on Copyright in the Digital Single Market created text and data mining exceptions. Scientific research gets a broad one. Commercial mining gets a narrower one that applies only where the rights holder has not opted out. The 2024 AI Act says the exception covers data collection for AI, and since August 2025 general purpose model providers have transparency obligations about their training data.

That framework makes a pure training claim hard to win. A defendant can point at the exception and argue the copying that happened during data collection was permitted, and then the fight moves to whether an opt out was properly expressed and honoured, which is a slow and technical argument. GEMA did not go that way. It argued that the model contains the lyrics, because it can emit them, and that a copy sitting inside a deployed product is a reproduction the exception does not cover.

That is a much cleaner claim. You do not need to litigate the crawl. You need a prompt, a screenshot, and the original lyrics side by side. The Munich court accepted that when the output matches the input almost word for word, the model has stored a copy, whatever the defendant calls the process that produces it.

The gap between training and retention

From a research standpoint the court drew a line that maps onto something real. Training on a text and memorising a text are different events. Most of what a language model sees during pretraining it never reproduces verbatim. A small fraction, weighted toward text that appears many times in the corpus, it can recite. Song lyrics are close to the worst case. They are short, they are duplicated across thousands of lyrics sites, and users ask for them by name.

So the legal exposure this ruling creates is concentrated in exactly the material that a model is most likely to have memorised, which is also the material rights holders can most easily test for. The practical consequence for a closed model provider is an output filter. Detect requests for lyrics, detect outputs that match known lyrics, refuse or truncate. That is roughly what the cease order demands, and it is feasible for a company that controls the serving stack.

What it implies for open models

The problem is that an output filter is a serving side control, and open weight models do not have a serving side. Once the weights are public, anyone can run them without the filter, and the memorised lyrics come out. If the court's logic is that a model containing recoverable copies is itself an infringing reproduction, then the infringement is in the file, and no downstream filter cures it.

That leaves an open model developer in Europe with two options that actually address the ruling. The first is to remove the memorised material before release, either by deduplicating and filtering the training corpus so that lyrics and similar high duplication text do not get memorised, or by post hoc unlearning, which is still unreliable. The second is to accept that the weights may contain recoverable copies and take the legal risk. The Munich court has just made the second option more expensive.

What we would like to see tested, by someone with the standing to publish it, is how much of the lyrics memorisation in current open models survives aggressive deduplication of the training set. If the answer is very little, then the fix is a data pipeline change rather than a legal one, and the open model community can get ahead of the next case rather than waiting for it.

Sources

  1. Wikipedia: Artificial intelligence and copyright (GEMA v. OpenAI and the EU text and data mining exception)