195 incidents on record · 2026 Headlights Incident reports by Ellie Harris · Melbourne
11 new this week Library last updated 12 September 2026
← The incident library
HD-INC-181
Probation and criminal justice · Netherlands · 2026 · Miscalibrated automated risk scoring used in probation advice, with the effect on individual advice not retrospectively determinable

A Dutch inspectorate found the probation organisations' recidivism tool had the formulas for detained and non-detained people the wrong way round, and had done since 2018

By Ellie Harris · Filed OXREC in use from 2018; the Inspectorate gave preliminary findings to the probation service and the ministry in mid 2025 and delivered its final report in December 2025

Alleged: Reclassering Nederland developed or deployed the AI system implicated in this incident. Details are drawn from public reports; parties are presumed innocent of any wrongdoing not established by an official finding.

A Dutch inspectorate found the probation organisations' recidivism tool had the formulas for detained and non-detained people the wrong way round, and had done since 2018

What happened

In February 2026, the Dutch Justice and Security Inspectorate reported serious problems with OXREC, a tool used by probation services to estimate whether someone was likely to reoffend.

OXREC had been used since 2018 and was run around 44,000 times a year. Its assessments helped inform advice given to prosecutors, courts and the prison service. But the Inspectorate found errors in how the tool had been set up. Most significantly, the formulas for people in detention and those outside detention had been swapped.

This was not a minor technical error. The Inspectorate estimated that around 21 per cent of general reoffending assessments could have ended up in a different risk category. The mistake pushed risk estimates too high for people who were not detained and too low for those who were.

There were broader concerns too. The Inspectorate found that the model relied on outdated data collected for a different purpose and population. It also raised concerns that factors such as neighbourhood and income could contribute to indirect discrimination.

The probation organisations stopped using OXREC on 12 February 2026.

What an auditable version would have shown

The state secretary reportedly told Parliament that it was impossible to determine whether the problems with OXREC had actually changed the advice given by probation officers, or whether that advice went on to affect decisions in the criminal justice system.

Probation officers did not rely on the tool alone. They used professional judgment alongside its risk assessment. But once the problems were discovered, there was no reliable way to go back and see how much influence the OXREC score had on an individual decision.

This is where a clearer record could have helped. For each assessment, it could have shown which version of OXREC was used, what risk category it produced, and how the probation officer considered that result when forming their advice. That would not replace human judgment or reduce it to a number. It would simply leave enough of a trail to understand what happened when something went wrong.

Where the gap was

It was reported that the Inspectorate found a risk that staff would too readily adopt the tool’s outputs and stop trusting their own judgment. It was reported that the Inspectorate recorded that staff had been told their own judgment was about as reliable as tossing a coin, which the Inspectorate said overstated the reliability of the algorithm.

A ConductRecord keeps the inputs, the model version, the score and what the person advising did with it. A MetricRecord is designed to count how often a score was followed and how often it was departed from. A ConstraintGate is designed to check a standing rule, such as whether a model has been validated for the population it is being run on, before the score is used.

None of these decides what risk category a person belongs in.

What governance should have looked like

It was reported that the probation organisations paused the tool on the day the report was published and said they would take up all of the Inspectorate’s recommendations. While OXREC is paused they said they would use RISC, a non-algorithmic instrument, to support structured professional judgment, and the government said a four eyes principle or collegial review could safeguard the quality of advice meanwhile.

Where an organisation uses an automated score to inform advice to a court, best practice would be to keep the score with the advice it informed, alongside the model version that produced it and what the person advising did with it. The organisation should also be able to say which population the model was validated on, and to check that before running it on anyone else.

Failure Pattern: it was reported that a risk prediction model used in probation advice ran for eight years with the formulas for detained and non-detained people swapped, and that the state secretary told Parliament the effect on individual advice could not be reconstructed afterwards.

Governance Principle: where an organisation uses an automated score to inform a decision about a person, it should be able to show which version of the model produced the score, what went into it and what the person advising did with it.

The reference implementation of ConstraintGate, ConductRecord and MetricRecord is open source. It lives at github.com/saffronandindia/headlights-oss, Apache 2.0 licensed and free to install. The repository is public now.

Sources

The mailing list

Fresh incident reports every week. One email to match.

We add new incidents to the library regularly, and send a single short email each week with what's new. The library stays free and open; this is just how you keep up with it.

No tracking. Unsubscribe in one click.

The record

An auditable system would have produced a signed, tamper-evident record the moment this happened: what the system did, the version that did it, the basis it acted on, and the action taken, and Reclassering Nederland could have produced it on demand.

This is the record the system as deployed did not produce in a signed, auditable form.

What this teaches
Capture what happened when it happens
What the system did, the version that did it, the basis it acted on, and the action taken, recorded at the moment, not reconstructed after.
Sign it, so no one has to trust the record-keeper
A tamper-evident entry. Edit it later and the signature breaks. The record does not ask for the benefit of the doubt.
Make it verifiable by anyone
A court, a regulator, a customer's lawyer can check the record themselves, without taking the company, or us, at our word.

Headlights summarises publicly reported AI incidents. All summaries are independently written, attributed to their original sources, and intended for research and educational purposes. Allegations are identified as such until established through official findings.

Last reviewed September 2026. This report is based on the sources listed above and reflects information available at the time of review; later developments may not be captured. Where a person is described as charged with or alleged to have done something, that allegation is unproven unless a conviction or a court or regulatory finding is stated. Headlights publishes journalism and commentary, not legal advice.

Want to write back?

Direct to my inbox.

ellie@useheadlights.com →