180 incidents on record · 2026 Headlights Incident reports by Ellie Harris · Melbourne
10 new this week Library last updated 30 August 2026
← The incident library
HD-INC-111
Public safety · Spain · 2025 · Under-rated automated risk assessment

Spain's VioGén algorithm rated a woman who reported her former partner as 'medium' risk, and she was killed three weeks later, in a case that fits a pattern an independent audit had already found

By Ellie Harris · Filed Report January 2025; death 9 February 2025

Alleged: Spanish Ministry of the Interior (VioGén) developed or deployed the AI system implicated in this incident. Details are drawn from public reports; parties are presumed innocent of any wrongdoing not established by an official finding.

Spain's VioGén algorithm rated a woman who reported her former partner as 'medium' risk, and she was killed three weeks later, in a case that fits a pattern an independent audit had already found

What happened

It was reported that in January 2025 a woman the BBC called Lina walked into a police station in Benalmádena, on Spain’s southern coast, to report that her former partner had threatened her. As with domestic-abuse reports across most of the country, the officers ran her case through VioGén, a system built by Spanish police and researchers that tries to gauge how likely a woman is to be attacked again. It asks 35 questions, about the violence and how bad it has been, the man’s access to weapons, his mental health, and whether she has left him or is about to, and turns the answers into a single rating: negligible, low, medium, high or extreme. Lina was rated medium, which brings a follow-up within 30 days rather than the escort or closer attention the higher levels can trigger. She asked a court in Málaga for a restraining order and was refused.

It was reported that about three weeks later, on 9 February 2025, she was dead. Her ex-partner, according to the BBC, let himself into her flat with his key, the home was soon on fire, and she did not get out. Her eleven-year-old son told police his father had killed her, and the man, the father of her three youngest children, was arrested. VioGén has been in use since 2007, and the rating it produces tends to hold: a 2014 study found officers accepted its assessment about 95 per cent of the time. How well those ratings stand up is hard to check from outside, because the system has not been opened to independent review. The one external audit on the record was done by the Eticas Foundation, which reconstructed how VioGén behaves from survivor interviews and public data after the authorities declined to take part. Eticas reported that between 2003 and 2021, 71 women later murdered by a partner or former partner had first been to the police about him, and that those logged in VioGén had been rated negligible or medium, not high. The interior ministry defends the system; its head of gender-violence research, Juan José López-Ossorio, has said that once a woman reports a man and is placed under police protection, the chance of further violence drops sharply. The ministry has said it will revise the questionnaire and abolish the lowest rating, negligible, and it has not released the case-by-case data that would let anyone outside measure how often the tool rates real danger too low.

What an auditable version would have shown

The ratings go in; what happens to the women afterwards is never tied back to them in any published figure. A system like this would put a person in the loop to check the rating before it sets how much protection someone receives, and would keep, for every assessment, the answers given, the rating produced and what that reviewer then decided, matched over time against what actually happened, so the share of medium and lower ratings that were followed by serious harm is a number the ministry itself can produce when asked. As it stands, the one version of that number in public was pieced together by an outside group working without the ministry’s help.

Where the gap was

The rating did a great deal of work in Lina’s case. It shaped the protection she was given, and a VioGén score is among the things a court weighs when it decides on a restraining order, though judges say it is only one of several factors they take into account. A 2014 study suggests officers seldom depart from it. Around a tool with that much influence, two things are missing. The first is any running measure of how often it is wrong. A MetricRecord supplies that: it tracks how each rating performs against what later happens to the women, so a tendency to rate danger too low shows up inside the system as a number rather than years later in an outside audit. The second is a real decision on top of the score rather than a rubber stamp. A ConductRecord supplies that: it keeps each assessment and the officer’s decision in enough detail to show a competent person looked at the case and decided, and could be asked why. None of this would, on its own, have saved Lina, and no record decides a case. But a rating that is followed nineteen times in twenty needs a person genuinely deciding behind it, and a way to catch it when it runs ahead of the facts.

What governance should have looked like

Where a rating sets how much protection a person in danger receives, two things have to be true. Someone accountable has to measure whether the rating is any good, by checking it against what later happens to real people and publishing the result, so the system is held to its own record rather than to reassurance about it. And the decision that follows the rating has to be a real one: an officer, or a court, weighing the case and able to say why they went where the tool pointed, or why they did not. VioGén is measured in public only through an audit its makers would not assist, and the data that would settle how often it rates danger too low has not been released.

If you or someone you know is affected by domestic or family violence, confidential support is available in Australia from 1800RESPECT on 1800 737 732, and in an emergency you can call 000.

Failure Pattern: a risk rating shaped the protection a person in danger received and was rarely overridden, while no published measure tracked how often the rating was set too low, so a pattern of under-rating surfaced only through an outside audit.

Governance Principle: where an automated rating determines the protection a person receives, its accuracy must be measured against real outcomes and published, and a competent person must make and record the decision that follows it, able to explain any reliance on the tool.

The reference implementation of MetricRecord and ConductRecord is open source. It lives at github.com/saffronandindia/headlights-oss, Apache 2.0 licensed and free to install. The repository is public now.

Sources

The mailing list

Fresh incident reports every week. One email to match.

We add new incidents to the library regularly, and send a single short email each week with what's new. The library stays free and open; this is just how you keep up with it.

No tracking. Unsubscribe in one click.

The record

An auditable system would have produced a signed, tamper-evident record the moment this happened: what the system did, the version that did it, the basis it acted on, and the action taken, and Spanish Ministry of the Interior (VioGén) could have produced it on demand.

This is the record the system as deployed did not produce in a signed, auditable form.

What this teaches
Capture what happened when it happens
What the system did, the version that did it, the basis it acted on, and the action taken, recorded at the moment, not reconstructed after.
Sign it, so no one has to trust the record-keeper
A tamper-evident entry. Edit it later and the signature breaks. The record does not ask for the benefit of the doubt.
Make it verifiable by anyone
A court, a regulator, a customer's lawyer can check the record themselves, without taking the company, or us, at our word.

Headlights summarises publicly reported AI incidents. All summaries are independently written, attributed to their original sources, and intended for research and educational purposes. Allegations are identified as such until established through official findings.

This report is based on the sources listed above and reflects information available at the time of review; later developments may not be captured. Where a person is described as charged with or alleged to have done something, that allegation is unproven unless a conviction or a court or regulatory finding is stated. Headlights publishes journalism and commentary, not legal advice.

Want to write back?

Direct to my inbox.

ellie@useheadlights.com →