195 incidents on record · 2026 Headlights Incident reports by Ellie Harris · Melbourne
11 new this week Library last updated 12 September 2026
← The incident library
HD-INC-188
Government and health technology · Canada · 2026 · A public body approved every supplier of a clinical drafting tool after its own procurement testing identified fabricated content and inaccuracies, with accuracy weighted at 4 per cent of the score and bias controls at 2 per cent

Ontario's auditor found that nine of twenty AI Scribe systems fabricated information in the government's own procurement testing, and that all twenty vendors were approved

By Ellie Harris · Filed Procurement conducted before the audit; audit published 12 May 2026

Alleged: Supply Ontario; OntarioMD; Ontario Health; Ministry of Health developed or deployed the AI system implicated in this incident. Details are drawn from public reports; parties are presumed innocent of any wrongdoing not established by an official finding.

Ontario's auditor found that nine of twenty AI Scribe systems fabricated information in the government's own procurement testing, and that all twenty vendors were approved

What happened

It was reported that on 12 May 2026 the Auditor General of Ontario published a performance audit of how the Ontario government uses artificial intelligence. It was one of four special reports released that day.

Part of the audit examined an AI Scribe request for bids process designed by Supply Ontario in consultation with OntarioMD, Ontario Health and the Ministry of Health. Twenty vendors were approved. An AI Scribe system turns a recorded consultation into a clinical note. In this testing the systems worked from recordings of two simulated consultations.

The audit found that nine of the twenty systems fabricated information and made suggestions about patients’ treatment plans. Evaluators saw notes stating that no masses were found, and notes recording an anxiety diagnosis, when nothing of the kind had been said in the recording. It found that twelve of the twenty produced notes recording a different drug from the one the doctor prescribed. It found that seventeen of the twenty missed key mental health details mentioned in the simulated conversations.

Vendors ran the test themselves. They were sent the two recordings, processed them offline, and submitted the notes their systems produced. Medical evaluators from OntarioMD and Ontario Health reviewed what came back. No live demonstration was required, and the audit records the risk that follows: a vendor could run a recording more than once, or alter the note. Vendors signed an attestation that the notes were unedited.

Accuracy carried a weighting of 4 per cent in the scoring. Bias controls carried 2 per cent. Security carried 11 per cent.

All twenty vendors were approved.

What an auditable version would have shown

The audit is the record. The counts of fabricated information, wrong drugs and missing mental health detail came from the evaluators’ review of the notes generated from the two simulated recordings. The audit also found that the procurement set no minimum passing scores for its evaluation criteria. Accuracy was tested, but no minimum accuracy score had to be met for approval. The audit recommended raising the weightings for security, privacy, bias and accuracy and assigning minimum passing scores.

A record of the evaluation would hold, for each vendor and each test recording, the audio submitted, the note returned, the discrepancies the evaluators marked, and the score that followed. Set that against the approval decision and it shows whether a fabrication cost a vendor anything.

The audit also found that eleven of the twenty approved vendors did not submit SOC reports, HITRUST certification or ISO 27001 certification, and that five did not submit a threat risk assessment or a privacy impact assessment. All vendors submitted declarations that they complied with the mandatory requirements and held the necessary documentation. On bias, the tender asked vendors to describe their organisational processes rather than to produce evidence such as testing results. The audit found no comprehensive evaluation of whether the systems mitigate the risk of unfair or biased outcomes.

Where the gap was

Fabricated content was identified, counted and written down. Accuracy carried a weighting of 4 per cent in the evaluation score. The gap, in this entry’s reading, was between what the buyer measured and what the measurement was allowed to decide.

A VerificationGate is designed to check a proposed output against a trusted source rather than against the model that produced it, here the drafted note against the recording it came from. A ConductRecord is designed to keep the recording, the model version, the note returned and the reviewer’s finding, so a claim about a system’s accuracy can be traced to the evidence behind it.

The audit’s recommendations do not include either design, although Recommendation 6 calls for Supply Ontario to require vendors to implement an IT control requiring users to attest that clinical staff have reviewed notes generated by AI Scribe systems. VerificationGate and ConductRecord are Headlights designs.

What governance should have looked like

The Auditor General recommended raising the weightings for security, privacy, bias and accuracy and setting minimum passing scores. It asked for a review of international guidance, including United Kingdom health service standards, and for vendors to implement an IT control requiring users to attest that clinical staff have reviewed the notes. It called for annual SOC 2 Type 2 reports from all vendors, with Supply Ontario and the Ministry of Health assessing whether those reports include the expected AI controls. It said Supply Ontario should require vendors to provide evidence of bias testing results, or conduct independent bias testing, before selecting vendors. And it recommended live demonstrations in future AI procurements.

Supply Ontario agreed with four of the five recommendations addressed to it. It partially agreed with the recommendation to raise the weightings and set minimum passing scores. The Auditor General has said it will follow up on implementation in two years.

Where a public body approves a tool that will write about patients, best practice would be to make the accuracy test capable of failing a supplier. That means a weighting that can change the result. It means a stated threshold below which a system is not approved. And it means a record of what each system was given, what it returned and what the evaluator found, in a form someone outside the procurement can check.

Failure Pattern: it was reported that a public body’s own procurement testing showed a drafting tool fabricating clinical content and recording the wrong drug, that accuracy carried a weighting of 4 per cent in the scoring and bias controls 2 per cent, and that all twenty vendors in the vendor of record arrangement were approved.

Governance Principle: where a public body approves a tool that will write about people, the evidence produced during the approval should be capable of failing a supplier, and the buyer should be able to show what it tested, what the tool returned and what it did about the difference.

The reference implementation of VerificationGate and ConductRecord is open source. It lives at github.com/saffronandindia/headlights-oss, Apache 2.0 licensed and free to install. The repository is public now.

Sources

The mailing list

Fresh incident reports every week. One email to match.

We add new incidents to the library regularly, and send a single short email each week with what's new. The library stays free and open; this is just how you keep up with it.

No tracking. Unsubscribe in one click.

The record

An auditable system would have produced a signed, tamper-evident record the moment this happened: what the system did, the version that did it, the basis it acted on, and the action taken, and Supply Ontario; OntarioMD; Ontario Health; Ministry of Health could have produced it on demand.

This is the record the system as deployed did not produce in a signed, auditable form.

What this teaches
Capture what happened when it happens
What the system did, the version that did it, the basis it acted on, and the action taken, recorded at the moment, not reconstructed after.
Sign it, so no one has to trust the record-keeper
A tamper-evident entry. Edit it later and the signature breaks. The record does not ask for the benefit of the doubt.
Make it verifiable by anyone
A court, a regulator, a customer's lawyer can check the record themselves, without taking the company, or us, at our word.

Headlights summarises publicly reported AI incidents. All summaries are independently written, attributed to their original sources, and intended for research and educational purposes. Allegations are identified as such until established through official findings.

Last reviewed September 2026. This report is based on the sources listed above and reflects information available at the time of review; later developments may not be captured. Where a person is described as charged with or alleged to have done something, that allegation is unproven unless a conviction or a court or regulatory finding is stated. Headlights publishes journalism and commentary, not legal advice.

Want to write back?

Direct to my inbox.

ellie@useheadlights.com →