Evidentiality Framework for AI
Download .md

Case study

The MAHA Report: Studies That Don’t Appear to Exist in a White House Health Report

In May 2025, a White House commission published a major report on children’s health. Some of the studies it listed as sources don’t appear to exist. Whether AI was used has never been confirmed.

What are the labels? The Evidentiality Framework asks an AI to mark each claim it writes: (g) generated, its own work; (u) given, passed to it by someone else; or (m) checked against a named source. The marks stay on a claim while people work with it. How the labels work.


What Happened


Follow the Claim

This is an illustration of how the labels would have worked, not a test. It only holds if the conditions under “What Would Have Had to Be True” held.

Follow one source: from a draft to a White House report

Key: as it happened, the type gets bigger as the claim sounds more certain. With labels: red (g) generated: written by the AI; green (u) given: passed on, with who said it; blue (m) checked against a named source.

As it happenedWith labels
Before 22 May 2025 · AI tool (suspected) → report writers
As it happened“Changes in mental health and substance abuse among US adolescents during the COVID-19 pandemic”, JAMA PediatricsA source that looks real, with a real scientist’s name on it.
With labels(g)Changes in mental health and substance abuse among US adolescents during the COVID-19 pandemic, JAMA Pediatrics(/g)If an AI tool wrote it and used the labels: marked as the AI’s.
Before 22 May 2025 · Draft → the list of sources
As it happenedListed as a source in the draft reportNow it looks like research someone read.
With labels(u)Changes in mental health and substance abuse among US adolescents during the COVID-19 pandemic, JAMA Pediatrics(/u: AI tool, unconfirmed)Passed on with where it came from, and that nobody has opened it.
Before 22 May 2025 · The check
As it happenedNobody looked it upThe source was never opened.
With labels(m)No such paper by the named author(/m: NOTUS, 29 May 2025, checked 3 Oct 2026)This check could be made at the time. We name the source that confirmed it later.
22 May 2025 · Report → the public
As it happenedPublished as a source in a federal reportNow it’s evidence in national health policy.
With labelsLeft out. The check found nothing, so the claim never reaches the finished document.Finished documents carry no labels. They only carry claims that passed the check.

The study does not appear to exist. Who wrote the report and which tools they used has not been made public; the first row shows how such a source would look if an AI tool produced it.


How the Labels Could Have Helped


What Would Have Had to Be True

The four conditions every case shares: the AI tool used the labels; it labelled its own work correctly (the weakest link: in our tests, AI sometimes mislabels its own work); the label stayed on when the text was copied; and someone owned a rule that unchecked claims don’t go further. In this case:


What Already Existed


What the Labels Wouldn’t Have Caught


Further Reading

Also listed in the AI Incident Database (#1084).

Next: All case studies · How the labels work