What happened
It was reported that Austria’s public employment service, the Arbeitsmarktservice or AMS, built an algorithm known as AMAS to help decide how to spend its limited budget for training and support. AMAS scored each jobseeker on their predicted chances in the labour market, drawing on a range of data about their work history and circumstances, and sorted them into three groups: people with good prospects who were thought to need little help, people in the middle who might benefit most from support, and people with poor prospects who were seen as hard to place. The middle group was steered towards the most help, so the score a person was given shaped what they were offered. According to an investigation by AlgorithmWatch in 2019, the model gave negative weight to being a woman, to being older, and to having a disability, and counted care obligations against women but not against men. A later academic study of the model found it also scored people who were not EU citizens lower.
It was reported that Austria’s data protection authority, the Datenschutzbehörde, examined AMAS and found it unlawful in 2020, stopping it before the nationwide rollout planned for the start of 2021. AMS challenged that decision. In September 2025 the Federal Administrative Court overturned the ban, finding that AMAS complied with the part of European data protection law that limits decisions made by a machine alone, because the caseworkers who used it kept a real say and could override the score. By then the system had been suspended for years and had never been rolled out across the country as AMS had planned. The court’s ruling settled the legal question of the caseworkers’ role. It did not settle whether the scores were fair to the people they counted against.
What an auditable version would have shown
The dispute was about who really made the decision, the caseworker or the score. From the outside, a caseworker who follows the score and one who weighs the case for themselves look the same. What tells them apart is a record: for each jobseeker, the score, whether the caseworker changed it, and why. The other thing no published figure showed is what the scoring actually did to the groups it counted against, how a woman’s score, or an older person’s, or a disabled person’s, compared with an otherwise similar man’s, and how that fed through to the support each was offered. Those are numbers a system can hold. Here they were argued over rather than produced.
Where the gap was
A score built partly on a person’s sex, age, disability and nationality helped decide what help they were offered, and the two questions it raised, whether a human really decided and whether the scoring disadvantaged those groups, were left to a regulator and a court rather than answered from records. A ConductRecord keeps each decision, the score and the caseworker’s part in it, in enough detail to show whether the human role was real, which is the point the court had to reason about from the system’s design. A MetricRecord keeps the aggregate: how scores, and the support that followed them, fell across sex, age, disability and nationality, so an effect on a protected group is a figure a regulator or a union can check. The court found the caseworkers’ role substantive. Whether the scoring treated those groups fairly was never measured in public.
What governance should have looked like
Where a score helps decide what public support a person receives, and is built partly on characteristics like sex, age and disability, two things should be in place. There should be a record showing that a person, not the score, made the decision, and able to show why. And the effect of the score on the groups it weights should be measured and published, so anyone can see how it changed the help those groups were offered. The court could say whether a real person made the decision, because the way the system was built showed that. Whether the scoring was fair to those groups, no one could say, because it was never measured.
Failure Pattern: a public-support score built partly on protected characteristics helped decide what help people were offered, and whether a human really decided, and whether the score disadvantaged those groups, were argued before regulators and courts rather than shown from records.
Governance Principle: where a score helps decide the public support a person receives, a human must make and record the decision, and the score’s effect across the groups it weights must be measured and published.
The reference implementation of ConductRecord and MetricRecord is open source. It lives at github.com/saffronandindia/headlights-oss, Apache 2.0 licensed and free to install. The repository is public now.
Sources
- Austria’s employment agency rolls out discriminatory algorithm (AlgorithmWatch)
- Algorithmic profiling of job seekers in Austria (Allhutter et al., Frontiers in Big Data, 2020)
- Austrian court (BVwG) rules the AMS public employment service algorithm complies with GDPR (Eurofound, EU agency)