> West Midlands Police: An AI-Invented Match in the Advice Behind a Fan Ban: Case study: a football match that never happened, later said to come from Copilot, was in police advice behind a fan ban. What labels could have done.
>
> Evidentiality Framework for AI. Early findings, October 2026. Web version: https://evidentiality-framework.org/case-west-midlands-police.html. Text CC BY 4.0.

Case study

# West Midlands Police: An AI-Invented Match in the Advice Behind a Fan Ban

In 2025, police advised a safety panel in Birmingham on whether a visiting football club’s fans could attend a match. Their advice included a match that never happened. The head of the force later said it came from an AI chatbot.

**What are the labels?** The Evidentiality Framework asks an AI to mark each claim it writes: (g) generated, its own work; (u) given, passed to it by someone else; or (m) checked against a named source. The marks stay on a claim while people work with it. [How the labels work](../../labels.html).

---

## What Happened

- Police were advising Birmingham’s Safety Advisory Group, a local panel that sets safety rules for big events, on whether Maccabi Tel Aviv fans could come to a match at Aston Villa on 6 November 2025. Police said the fans were high-risk, citing earlier trouble around a match in Amsterdam.
- On 10 October 2025, the senior officer in charge wrote to the panel’s chair. The letter said Maccabi Tel Aviv had last played in the UK against West Ham. No such match took place.
- On 16 October the panel decided visiting fans should be banned, after spoken briefings from police that didn’t repeat the claim. On 24 October it looked at the question again from scratch. The police’s written report for that meeting said: “The most recent match Maccabi Tel Aviv played in the UK was against West Ham United … on 9th November 2023.” The ban stayed, and the match was played without away fans.
- On 6 January 2026, the head of West Midlands Police (the chief constable) told MPs, the members of Parliament looking into it, that the force did not use AI. On 12 January he wrote to correct this: the claim came from Microsoft Copilot. He retired on 16 January, after the Home Secretary, the minister in charge of policing, said she had lost confidence in him.
- In February 2026 a committee of MPs found that the force “failed to do even basic due diligence”, and that some of its key claims about the Amsterdam trouble “originated from a query to Microsoft Copilot AI”. It also found the chief constable “did not intentionally mislead” them.
- The police watchdog, the Independent Office for Police Conduct, opened an investigation in January 2026. In August it told the former chief and four others that their conduct is being investigated. That is not a finding that anyone did wrong.

---

## Follow the Claim

This is an **illustration** of how the labels would have worked, not a test. It only holds if the conditions under “What Would Have Had to Be True” held.

Follow the West Ham match: one AI answer, from a chatbot to a fan ban

**Key:** as it happened, the type gets bigger as the claim sounds more certain. With labels: red (g) generated: written by the AI; green (u) given: passed on, with who said it; blue (m) checked against a named source.

**As it happened****With labels**

Before 10 October 2025 · Copilot → an officer

As it happenedA match against West Ham, with a dateAn AI answer, stated as fact. (The chief later said Copilot; the inspector heard conflicting accounts.)

With labels(g)Maccabi Tel Aviv last played in the UK against West Ham United, on 9 November 2023(/g)If the tool used the labels: marked as written by the AI.

10 October 2025 · Senior officer → panel chair

As it happenedIn a letter to the safety panelNow it sounds like police intelligence.

With labels(u)Maccabi Tel Aviv last played in the UK against West Ham United, on 9 November 2023(/u: Copilot, unconfirmed)Passed on as something the officer was given, with who said it.

Before 24 October 2025 · The check

As it happenedNobody checkedThe claim was never tested.

With labels(m)No such match took place(/m: HM Chief Inspector of Constabulary, letter of 14 January 2026, checked 3 Oct 2026)Football fixture lists could show this at the time. We name the source that confirmed it later.

24 October 2025 · Police report → safety panel

As it happenedRepeated in the written report; the ban staysNow it’s part of the case for a decision.

With labelsLeft out. The check found nothing, so the claim never reaches the finished document.The panel would still decide; this one false claim wouldn’t be part of it.

No match between West Ham and Maccabi Tel Aviv took place. The quote is the wording in the police report, as given by HM Chief Inspector of Constabulary.

---

## How the Labels Could Have Helped

- **The guess stays marked as a guess.** The claim keeps a label saying it came from Copilot and nobody confirmed it, so it can’t pass as police intelligence.
- **The report checks before it repeats.** A claim still marked unconfirmed gets checked before it goes into advice to the panel.
- **Where it came from is written down.** When MPs asked, the answer would have been on the working notes.

---

## What Would Have Had to Be True

The [four conditions every case shares](../../cases.html#assume): the AI tool used the labels; it labelled its own work correctly (the weakest link: in [our tests](../../check.html), AI sometimes mislabels its own work); the label stayed on when the text was copied; and someone owned a rule that unchecked claims don’t go further. In this case:

- The rule is owned by the senior officers preparing the advice. Advice to a panel is “finished” work, so labels stay in the working notes and only checked claims go in. *Open question: internal decision documents may be better treated as working documents that keep their labels.*

---

## What Already Existed

- **Police already grade intelligence.** The UK’s 3×5×2 system grades the source from 1 (reliable) to 3 (not reliable), and the information from A (known directly) to E (suspected false). A chatbot answer would most likely be 2 (untested) and D (“not known”).
- **This claim went around it.** The Chief Inspector found that not all of the report went through the force’s intelligence unit.
- **A simpler check would have caught it:** a fixture list. Labels add one thing: every unchecked claim is marked, not only the ones someone thinks to look up.

---

## What the Labels Wouldn’t Have Caught

- **Skipping the process.** The claim went around the force’s own grading. A label can be skipped the same way.
- **The other failings.** MPs found the force relied on disputed claims about the Amsterdam trouble, didn’t consult the local Jewish community, and wrongly told the panel it had. Labels on one AI line fix none of that.
- **Lost records.** The officer who attended a meeting with Dutch police threw away his handwritten notes of it, the Chief Inspector found. No label helps with notes that no longer exist.

---

## Further Reading

- [Home Affairs Committee report, HC 1553 (February 2026)](https://committees.parliament.uk/publications/51721/documents/286921/default/)
- [HM Chief Inspector of Constabulary: letter to the Home Secretary (14 January 2026, PDF)](https://data.parliament.uk/DepositedPapers/Files/DEP2026-0017/Letter_from_HMICFRS_to_Home_Sec_re_WMP.pdf)
- [Independent Office for Police Conduct: investigation update](https://www.policeconduct.gov.uk/node/13386)
- [College of Policing: how intelligence is graded (3×5×2)](https://www.college.police.uk/app/intelligence-management/completing-intelligence-report)
- [The Register: police chief retires over AI hallucination](https://www.theregister.com/2026/01/19/copper_chief_cops_it_after/)
- [ITV: watchdog investigates former chief constable](https://www.itv.com/news/central/2026-08-26/former-chief-constable-probed-over-maccabi-tel-aviv-supporter-ban)

Also listed in the [AI Incident Database (#1400)](https://incidentdatabase.ai/cite/1400/).

**Next:** [All case studies](../../cases.html) · [How the labels work](../../labels.html)
