Automating with AI? Australia cut one human check from its welfare debt system and unlawfully billed 433,000 people
Jamie Watters
Operational resilience and AI delivery practitioner.

On 11 July 2016, Australia's welfare department switched on a new online system: the next stage of a debt-recovery scheme it had been piloting since 2015. About a week later, one of the people running it emailed colleagues with the subject line "OCI – Our future is here". The first case had gone through.
The scheme became known as Robodebt. In 2021 the Federal Court approved a class-action settlement after the government admitted it had no proper legal basis for debts based on averaging people's income. The evidence, the court said, showed debts of at least $1.763 billion unlawfully asserted against approximately 433,000 Australians (Prygodicz v Commonwealth of Australia (No 2)).
It wasn't AI. But it's a fully documented case of a mistake anyone automating work can make. You cut the slow human step to go faster, and find out too late that it was the step asking whether the data was right.
The slow way
The department already matched what employers told the tax office against what people on benefits had declared. Only the riskiest got a person: about 20,000 files a year, picked for the strongest likelihood of an overpayment (Royal Commission report, 2023).
Each one went to a compliance officer, who checked the file before anyone contacted the person. The Royal Commission's report gives an example. Someone declares they work for Big W. The tax data says Woolworths. To a computer, that's a mismatch. To an officer who knows Big W's legal name is Woolworths, it's nothing. They update the record and move on.
If they couldn't explain it, they asked the person, then the employer for payslips. It was slow. It was also the part of the process that knew things the data didn't.
The fast way
The new system kept the matching and dropped the officer's check. From then on, the Commission found, "there was no intervention by a compliance officer prior to the discrepancy being raised with the recipient."
A letter went out instead. If the person didn't answer, accepted the tax office figure, or couldn't prove otherwise, the system took the employer's figure and "divided up evenly" across the fortnights in the period. Any fortnight that came out too high became a debt, sometimes with a 10 per cent penalty added automatically, the Commission found.
That works if your pay is the same every fortnight. If it goes up and down, the Commission says, averaging was "most unlikely to produce an accurate debt amount". And it says the vast majority of recipients had pay that went up and down.
The Commission lists what changed. First and biggest, the tax office figure was treated as the truth, unconfirmed by anyone. Second, the officer's checks were swapped for a letter that put the burden on the person. Put the two together and nothing in the process asked whether the figure was right.
The minister's media release, quoted in the Commission's report, said the system was "now initiating 20,000 compliance interventions a week", up from 20,000 a year.
The check that came back bad
They did check it. From 15 July, while numbers were small, each assessment was held back from the person until it had been checked. Compliance officers redid the computer's debts by hand. By 30 September they had done 919 of these checks, and marked 29 per cent of the computer's debts as wrong.
That's the Commission's count, and it thinks the real figure was higher. The officers probably redid the sums with the same averaging the computer used. A check that uses the rule to test the rule won't catch the rule being wrong.
The same check showed, on the Commission's figures, that 70 per cent of people sent a letter didn't respond. So the averaging wasn't a last resort for a few. It was what happened to most people.
It wasn't a formal trial with a pass mark. The Commission says it was rolled out in full with "no proper evaluation" of the pilots first. But the problem showed up in the daily and weekly reports, and the Commission called the results "extremely concerning".
The volumes went up anyway. On 20 September, under pressure to hit the budget numbers, a manager set out a plan to start 63,000 extra reviews in the week of 26 September. The check ran until 30 September. A staff member proposed adding the risks to a briefing paper. She told the Commission the reply was: "under no circumstances are you to put this in writing."

The scale-up started four days before the check finished.
The bill
On the Federal Court's figures, the government recovered about $751 million from about 381,000 people. The Royal Commission put the scheme's net cost to the government at about $565 million. The budget measure behind it had promised savings of $1.7 billion over five years. The Commission called Robodebt "a crude and cruel mechanism, neither fair nor legal".
Two limits. The Commission is plain that "the Scheme did not use AI": it was "a system of business rules". And the scale-up wasn't an accident. People inside raised the problems, and the volume targets came first.
The slow step is often the control
When you map a process to automate it, the slow human steps stand out. They look like waste, and sometimes they are. But sometimes the slow step is where someone knows Big W is Woolworths, or asks whether the figure is right at all. Cut it, and the rule runs alone on every case it gets wrong, at full speed.
I think AI makes that cut more tempting, because the agent is fast and the person is slow.
There's a second version of the same trap: the step that matters isn't in the documents. Vasuman Moza, who builds AI agents for large companies at Varick Agents, gave this example in a talk at the AI Engineer conference:
"Sarah in AP handles the workflow today in this way, but when things go wrong, she actually sends it over to Chris, who then takes four days of cycle time to handle reconciliations between a purchase order and an invoice."
Company documents, he says, cover "the golden path and maybe an edge case or two". Sarah and Chris are an unnamed example, not an identified client, and his firm starts every job with an audit, so he has a reason to say it. Robodebt's officer was on the process chart, and someone chose to cut that step. If Chris isn't on the chart, an agent built from the chart skips him without anyone deciding to.

The agent is built from the top line. The cost sits in the bottom one. My illustration, built on the example in Moza's talk.
My career has been business continuity, disaster recovery and crisis management. In crisis exercises it's not uncommon for people to ad-lib, to not know their own procedures, or to miss cues they should respond to. The written plan doesn't show you that. The exercise is how you find it. An automation that reads the document, builds the golden path and goes live is running its first exercise in production.
Five things to do before AI takes over a process
- Before you cut a slow step, ask what it caught last time. Get a real case with a date, then ask how often it happens and what it cost. If it caught something the rule would have missed, treat it as a control until you've shown otherwise.
- Find the steps that aren't written down. Ask each person who they send the odd cases to, then ask that person the same question. Look for the side channels: email, chat, a phone call, a personal spreadsheet.
- Run the new process alongside the people on real cases. Check the hard ones against the evidence the old step used: the payslip, the file, a phone call. Agreeing with the person isn't enough. Robodebt's officers checked the computer with the computer's own method, and the Commission thinks that's why the count came out low.
- Decide before the test what result stops the launch. Write it down, and name who can stop it: someone who isn't chasing the target. Make some mistakes an instant stop, like one payment to the wrong supplier. Robodebt's check came back 29 per cent wrong, and a staff member says she was told not to put the risks on paper. A number only stops a launch if people agreed beforehand that it would.
- After go-live, check a sample of the cases it handled on its own, with no person involved. Robodebt's wrong debts were exactly those: the rule saw nothing odd, so no officer looked. Complaints will tell you eventually. A sample tells you sooner.
These won't tell you whether the rule itself is sound. Robodebt's biggest problem was the averaging, and mapping the process doesn't fix a rule that's wrong. What they tell you is what you're about to lose.
One step, worked through. The example and the numbers are made up.
| Example: match invoice to purchase order | |
|---|---|
| What it caught last time | An invoice above the order price. The supplier had raised prices and the order hadn't been updated |
| How often | Say 9 of the last 100 invoices |
| Who handles it now | Chris, sent over by Sarah by email |
| What Chris checks that the rule doesn't | The supplier's price letter, and a call to the buyer |
| What the AI hands over | Both documents, the difference, and what it has already checked |
| Mode, and why | Human in the loop: Chris approves before payment and the AI can't release a payment itself, because a wrong payment is hard to get back |
| Can Chris keep up? | At 400 invoices a day, 9 in 100 is 36 a day for Chris. If he can't clear that, fix the queue first |
| What stops the launch | Say: any payment to the wrong supplier, or Chris overruling the AI on more than 1 in 10 of the mismatched invoices in the test |
| Who can stop it | The finance manager, not the project lead |
Process mining software can help with step 2, but only with what your systems record. If Sarah forwards the case from her inbox and nothing links the email to the invoice, the tool sees an invoice that sat still for four days. Ask Sarah why.
Robodebt had the slow step: an officer who knew Big W was Woolworths. It had a check, and the check marked 29 per cent of the debts wrong. It turned the volume up anyway.
Sources
- Royal Commission into the Robodebt Scheme, Report, July 2023 (accessible full report, read from the Internet Archive copy of the Commission's own file): pp. xxiv, xxvi, xxix, 131 to 133, 401, 471 to 475.
- Prygodicz v Commonwealth of Australia (No 2) [2021] FCA 634, Federal Court of Australia, 11 June 2021, paras 4, 5 and 8.
- Vasuman Moza, Varick Agents, "AI tools for Forward Deployed Engineering", AI Engineer conference talk.