Mata v. Avianca: the brief with six invented cases
Judge Castel sanctioned two New York lawyers and their firm on June 22 for filing opinions ChatGPT had fabricated and then standing by them for three months. What the opinion actually found, why asking the model to verify itself was the fatal step, and what it means for grounded tools.
What happened
Roberto Mata sued Avianca over a serving cart that hit his knee on a flight from El Salvador to JFK. The airline moved to dismiss under the Montreal Convention's limitations period. On March 1, 2023, Peter LoDuca of Levidow, Levidow and Oberman filed an affirmation in opposition, written by his colleague Steven Schwartz, that cited and quoted from decisions in the Federal Reporter, the Federal Supplement and Westlaw. Avianca's reply on March 15 said its lawyers could not find most of the cases and set the ones they could not locate apart in quotation marks.
Judge P. Kevin Castel's opinion of June 22 lays out what followed. The court could not find the cases either and on April 11 and 12 ordered LoDuca to file copies. What came back on April 25 were excerpts, not full opinions, attributed to an unnamed online database, plus a note that one cited case, Zicherman, could not be found at all. The excerpts were fabrications. The Eleventh Circuit's clerk confirmed that Varghese v. China Southern Airlines did not exist and that no party by that name had appeared before the court since 2010. The panel listed for it included a judge who sits on the Fifth Circuit.
The step that did the damage
Schwartz had used ChatGPT to research a federal bankruptcy question he said he was completely unfamiliar with. That alone is not what the court sanctioned. The opinion opens by saying there is nothing improper about using a reliable AI tool for assistance, and that the rules put a gatekeeping duty on the lawyer to check what goes in a filing. The sanction was for what happened after the doubts arrived.
The record shows Schwartz entered the Varghese citation into a free case lookup site and could not find it, before or shortly after filing. His answer under oath was that he could not conceive of the tool fabricating a case and assumed it was unpublished or hard to access. When the court's order to show cause arrived, he did the thing that still gets people into trouble. He asked ChatGPT whether Varghese was real and whether the other cases were fake. The model said they were real and could be found on Westlaw and LexisNexis. Those screenshots became Appendix B of the sanctions order.
The court treats this as the centre of the case. Schwartz's May 25 affidavit presented the screenshots as evidence that the tool had assured him of its reliability. His June 6 declaration reframed the same exchange as confirming his suspicion that the model was answering without regard for truth. Judge Castel called these shifting and contradictory explanations and used them to find subjective bad faith. The lesson for anyone building professional tools is that a model's own claim about its sources carries no evidential weight, and that a workflow which lets the user ask the model to verify itself is a workflow with no verification step.
What the court actually sanctioned
The findings are specific and worth listing because the popular version of the story is just that a lawyer used ChatGPT. LoDuca signed the March 1 affirmation under penalty of perjury without reading a single cited case, and swore to the April 25 affidavit after Schwartz walked it into his office twenty feet away. LoDuca also obtained an extension by telling the court he was on vacation when he was not. Schwartz filed excerpts of cases he already knew he could not find in full, and the court found the Varghese text to be gibberish on its face, with a procedural history that borders on nonsensical and two different plaintiffs.
The penalty was 5,000 dollars, paid jointly into the court registry, plus letters. Within 14 days the respondents had to write to Mata attaching the opinion and the hearing transcript, and to each of the six judges falsely named as authors of the fake Varghese, Shaboon, Petersen, Martinez, Durden and Miller opinions. The court said it chose the amount to deter, not to compensate, and said explicitly that the outcome would have been different had the lawyers come clean after the March 15 reply.
What this does to professional tooling
The opinion is a list of harms from a fake citation. The opponent wastes money exposing it. The court's time is lost. The client loses real arguments. Judges get named as authors of things they never wrote. And a future litigant may claim doubt about a genuine ruling. Each of those maps to a requirement a legal research product now has to meet. Every citation must resolve to a document the product actually holds, every quote must be found in that document, and the tool must be able to say, in a way the user cannot argue with, that a citation does not exist.
That is a retrieval architecture, and the case makes the argument for it better than any benchmark. A model answering from parametric memory produced citations with correct-looking reporters, plausible docket numbers and real judges' names. Only the existence check failed, and the existence check is the one thing a database does perfectly. Products that ground generation in a corpus and refuse to emit a citation that is not in it would have made this filing impossible.
What we take from it
The lawyer's error was one that researchers make too. He treated fluency as evidence and treated the model's confidence in its own output as a second source. Both of those are habits worth naming in any group that uses these tools to write, because a model that can fabricate a case can fabricate a citation to a paper, and the check is the same. Look it up somewhere the model did not write.
We would like to see the vendors of legal tools publish their hallucinated citation rates on a held-out set of real briefs, with the retrieval layer on and off. The Avianca facts give the field a clean test case, a bankruptcy tolling question under the Montreal Convention, and a court has now written down what a wrong answer costs.
Sources
From the foundation