What happened
It was reported that Australia’s Merit Protection Commissioner, the office that checks whether public-service jobs are filled on merit, raised concerns about an AI-assisted bulk recruitment round. In its 2021-22 review the Commissioner overturned twelve promotion decisions for the year, eleven of them from that single bulk round, and warned about the risks these tools pose to hiring on merit. About a quarter of the agencies surveyed had used AI-assisted or automated recruitment tools in the previous year.
It was reported that the concern resurfaced in December 2024, when the Community and Public Sector Union gave evidence to a parliamentary committee examining the public sector’s use of artificial intelligence. The union’s deputy secretary, Rebekah Fawcett, was reported to have told the committee the union was aware of cases “in Services Australia, workers with a proven track record being rated unsuitable and ruled out for promotion or permanency because their video recording or written application failed to use language that the algorithm was searching for.” She said the union was also aware, anecdotally, of tools appearing to favour candidates who had used generative AI to write their applications. A union survey of about 1,800 Commonwealth staff was reported to have found most were concerned about AI in recruitment and promotion, and the union’s national secretary called for AI to be barred from making recruitment decisions.
Services Australia disputed the union’s account. It was reported that a spokesperson said all recruitment decisions, including whether an applicant is suitable, are made by a delegated staff member in line with public-service merit principles, and that the agency does not use AI-assisted tools to screen resumes or assess video interviews. The spokesperson said the agency does use AI-assisted tools, through service providers, to help assess applications, that going without them would make recruitment slower and more expensive, and that the tools are trained and quality-checked by staff to reduce bias. The two accounts differ on a central point: the union said staff were rated unsuitable by an algorithm, while the agency said a delegated staff member, not an algorithm, makes each decision. As of mid-2026, no regulator, tribunal or court had been reported to have ruled on the specific claims.
What an auditable version would have shown
This is really an argument a record could settle. If experienced staff were screened out, the questions are basic ones: what did the tool score them, what did it weigh, and did a person actually make the call or just wave it through? Each of those has an answer, if the process keeps the right records. For each application there would be a signed record of what the tool put out and why, and whether a delegated decision-maker genuinely looked at the result before signing it. Put those together and you can also see how often the tool’s rankings are overturned later, the very thing the Merit Protection Commissioner had to work out case by case. With a record like that, “a human decided” and “the algorithm decided” stop being rival claims, because one of them is on file.
Where the gap was
An applicant could be knocked back, and neither they nor, on the union’s account, anyone else could say why. A ConductRecord holds the basis of each screening decision, what the tool produced, what it weighed, and who signed off, so a candidate can be told the reason and a reviewer can check it. A MetricRecord keeps the tool’s outcomes as a number worth watching, the share of its rankings that get overturned on review, so a run of contested or overturned decisions shows up as a figure rather than trickling in one appeal at a time. The Merit Protection Commissioner found the bad promotion decisions the hard way, one at a time, while a standing record is how you would catch them sooner.
What governance should have looked like
Public-service hiring runs on a promise that the job goes to the best person on the merits, and a tool that helps sort applicants has to be held to that promise, whoever makes the final call. In practice that means each decision it touches should be something you can explain to the person affected, and the process should keep track of how often its picks are overturned, especially where strong candidates get missed because their wording did not match what the tool was scanning for. The thread running from the Merit Protection Commissioner’s findings to the union’s evidence is a simple one: trusting a hiring tool is not the same as showing it works. An agency should be able to show, from its own records, both who made each call and how the tool has actually held up when its decisions were reviewed.
Failure Pattern: job applicants were filtered with the help of an algorithm whose basis they could not see, and there was no published measure of how often its outcomes were overturned.
Governance Principle: when a tool helps decide who is hired, each decision should carry a record of the basis it rested on, and the process should track how often its outcomes are overturned, so error and bias can be seen rather than assumed away.
The reference implementation of ConductRecord and MetricRecord is open source. It lives at github.com/saffronandindia/headlights-oss, Apache 2.0 licensed and free to install. The repository is public now.
Sources
- Public sector AI hiring practices under scrutiny in APS (The Canberra Times, 9 December 2024)
- AI recruitment threatens merit principle, union warns (The Mandarin, 5 December 2024)
- Inquiry into the use and governance of AI systems by public sector entities (Joint Committee of Public Accounts and Audit)