Case study
Mata v. Avianca: Court Cases Invented by ChatGPT, Filed in Federal Court
In 2023, two New York lawyers cited earlier court cases in a lawsuit. The cases didn’t exist. ChatGPT had made them up, and when asked, it said they were real.
What are the labels? The Evidentiality Framework asks an AI to mark each claim it writes: (g) generated, its own work; (u) given, passed to it by someone else; or (m) checked against a named source. The marks stay on a claim while people work with it. How the labels work.
What Happened
- A man sued the airline Avianca. The case was moved to a New York federal court. On 1 March 2023 his lawyers filed a brief, a written argument to the court, citing earlier cases, including “Varghese v. China Southern Airlines Co., Ltd., 925 F.3d 1339 (11th Cir. 2019)”. That case doesn’t exist.
- On 15 March, Avianca’s lawyers told the court they couldn’t find several of the cases. The court ordered copies. In April the lawyers filed what they said were excerpts.
- The lawyer who did the research had asked ChatGPT “Is Varghese a real case”. It told him the cases were real and could be found in the main legal databases. The lawyer who signed the filing hadn’t read any of the cases.
- On 22 June 2023 the judge fined the two lawyers and their firm $5,000, and ordered them to send the court’s ruling to the real judges falsely named as authors of the fake rulings.
- The judge wrote that “there is nothing inherently improper about using a reliable artificial intelligence tool for assistance.” He found bad faith, based on “acts of conscious avoidance and false and misleading statements to the Court”.
Follow the Claim
This is an illustration of how the labels would have worked, not a test. It only holds if the conditions under “What Would Have Had to Be True” held.
Key: as it happened, the type gets bigger as the claim sounds more certain. With labels: red (g) generated: written by the AI; blue (m) checked against a named source.
The case does not exist. The reference is as filed, quoted in the court’s ruling.
Key: as it happened, the type gets bigger as the claim sounds more certain. With labels: red (g) generated: written by the AI.
The court found the lawyer’s explanations of when he asked this were not consistent.
How the Labels Could Have Helped
- Self-checks stay red. Asking the AI whether its answer is real produces another (g), not an (m). This is the clearest lesson of the case.
- Filing waits for a check. A case nobody has found in a legal database isn’t cited.
- The signing lawyer can see what was checked. He signed without reading the cases. Labels would have shown him that nobody had.
What Would Have Had to Be True
The four conditions every case shares: the AI tool used the labels; it labelled its own work correctly (the weakest link: in our tests, AI sometimes mislabels its own work); the label stayed on when the text was copied; and someone owned a rule that unchecked claims don’t go further. In this case:
- The rule is owned by the lawyer who signs the filing, who already has that duty.
What Already Existed
- Lawyers already have to check. US court rules (Rule 11) make a lawyer who signs a filing vouch that its legal arguments rest on real law.
- Tools exist. Legal databases and “citators” confirm whether a case exists. The firm’s research service had limited coverage of federal cases, the court found.
- A simpler check would have caught it: a search in a legal database. Labels add one thing: the unchecked references stand out before anyone signs.
What the Labels Wouldn’t Have Caught
- What came after. Much of the court’s criticism was about how the lawyers responded once they were warned. Labels don’t make anyone own up.
- Signing without reading. A label only helps if the person signing reads it.
- Limited research tools. The firm turned to ChatGPT partly because its usual service didn’t cover the cases it needed.
Further Reading
- Mata v. Avianca, the court’s ruling on sanctions, 22 June 2023 (Justia)
- Bloomberg Law: phony ChatGPT brief leads to $5,000 fine
Also listed in the AI Incident Database (#541).
Next: All case studies · How the labels work