180 incidents on record · 2026 Headlights Incident reports by Ellie Harris · Melbourne
10 new this week Library last updated 30 August 2026
← The incident library
HD-INC-130
Education technology · United States · 2021 · Biased face-detection and behaviour flagging in exam surveillance

Universities and exam boards watched students through remote-proctoring AI that flagged them for cheating, and its face-detection worked so poorly on darker-skinned students that some could not start their exams

By Ellie Harris · Filed Widely used during 2020 and 2021; disparities reported 2020 to 2022

Alleged: Proctorio; ExamSoft; and other remote-proctoring vendors developed or deployed the AI system implicated in this incident. Details are drawn from public reports; parties are presumed innocent of any wrongdoing not established by an official finding.

Universities and exam boards watched students through remote-proctoring AI that flagged them for cheating, and its face-detection worked so poorly on darker-skinned students that some could not start their exams

What happened

It was reported that when exams moved online during the pandemic, universities and exam boards turned to remote-proctoring software from companies such as Proctorio and ExamSoft. The software watched a student through their webcam while they sat an exam and used AI to detect the student’s face, follow their eyes and movements, and flag anything it judged to be a sign of cheating, such as looking away or another face appearing. A flag could then be reviewed, but the student was being judged, throughout, by a system reading them through a camera.

It was reported that the face-detection worked poorly on darker-skinned students. An analysis of the face-detection model used by Proctorio found that it failed to detect Black faces more than half the time, and a peer-reviewed study found disparities in how the software performed across skin tone, race and sex. Students described having to light their faces harshly or being unable to get the software to recognise them at all, which could stop them from starting the exam. In 2020 a group of US senators wrote to the proctoring companies raising concerns about bias and privacy. A student the software flagged could not see what had triggered it, and the companies defended their tools.

What an auditable version would have shown

Software that decides a student’s face cannot be seen, or that their behaviour looks like cheating, is making judgements the student cannot see and cannot answer. An auditable version keeps a record for each student showing what the system detected, what it flagged and why, and makes it available, so a student wrongly flagged, or wrongly told their face cannot be found, has something concrete to point to. It would also count how often the software fails to find a face or raises a flag, broken down by skin tone, so a tool that works worse for darker-skinned students is something the school and the vendor can see, not a pattern students have to prove one complaint at a time.

Where the gap was

Students were judged by software reading them through a camera, and its face-detection failed on darker-skinned students while its flags could not be seen or contested. A MetricRecord counts how often the system fails to find a face or raises a flag across skin tones, so a tool that performs worse for some students is measured before it is relied on, not discovered by the students it fails. A ConductRecord keeps each flag and what triggered it, and puts it in the student’s hands, so a wrong flag can be shown and put right. The disparity here was found by an outside analysis of the model. A record kept inside the software would have shown the same thing from the start.

What governance should have looked like

Where software judges people through a camera, it should be tested and shown to work across skin tones before anyone is made to rely on it, and each judgement it makes should be recorded and open to the person to see and challenge. Best practice would be for the vendor and the school to measure and publish how often the system fails to detect a face or raises a flag across different groups of students, and to keep, for each student, what was flagged and why, available to them. The companies said their tools worked. Whether they worked for a darker-skinned student trying to start an exam is something a measurement across skin tones would show.

Failure Pattern: exam-surveillance software judged students through a webcam and its face-detection failed on darker-skinned students, locking some out or flagging them, while judging behaviour the student could not see or contest.

Governance Principle: where software judges people through a camera, its detection must be tested and shown to work across skin tones before it is relied on, and each flag and its basis must be recorded and open to the person to see and challenge.

The reference implementation of MetricRecord and ConductRecord is open source. It lives at github.com/saffronandindia/headlights-oss, Apache 2.0 licensed and free to install. The repository is public now.

Sources

The mailing list

Fresh incident reports every week. One email to match.

We add new incidents to the library regularly, and send a single short email each week with what's new. The library stays free and open; this is just how you keep up with it.

No tracking. Unsubscribe in one click.

The record

An auditable system would have produced a signed, tamper-evident record the moment this happened: what the system did, the version that did it, the basis it acted on, and the action taken, and Proctorio; ExamSoft; and other remote-proctoring vendors could have produced it on demand.

This is the record the system as deployed did not produce in a signed, auditable form.

What this teaches
Capture what happened when it happens
What the system did, the version that did it, the basis it acted on, and the action taken, recorded at the moment, not reconstructed after.
Sign it, so no one has to trust the record-keeper
A tamper-evident entry. Edit it later and the signature breaks. The record does not ask for the benefit of the doubt.
Make it verifiable by anyone
A court, a regulator, a customer's lawyer can check the record themselves, without taking the company, or us, at our word.

Headlights summarises publicly reported AI incidents. All summaries are independently written, attributed to their original sources, and intended for research and educational purposes. Allegations are identified as such until established through official findings.

This report is based on the sources listed above and reflects information available at the time of review; later developments may not be captured. Where a person is described as charged with or alleged to have done something, that allegation is unproven unless a conviction or a court or regulatory finding is stated. Headlights publishes journalism and commentary, not legal advice.

Want to write back?

Direct to my inbox.

ellie@useheadlights.com →