Live data from Hacker News

We investigated Amsterdam's attempt to build a 'fair' fraud detection model

lighthousereports.com

41–50 of 87 posts

Re: We investigated Amsterdam's attempt to build a 'fair' fraud detection model

#41

[flagged]

> Why would you assume that all groupings of people commit welfare fraud at the same rate? What's the alternative? It's an unattainable statistic, the people who get away with crime. Instead, what ends up getting used is the fraud rates under the old system, or ad hoc rules of thumb based in bigoted anecdotes. So instead you delcare that you don't think that ethnicity is in and of itself a cause of fraud. Even if the…

One of these days, I’m still hopeful, we will figure out that behaviors are taught, usually by parents. Intentionally or accidentally, kids learn from what they see.

I don’t care what nationality you are, or what your skin color happens to be, the root cause is how kids are reared.

In my head it’s so simple.

Re: We investigated Amsterdam's attempt to build a 'fair' fraud detection model

#42

> But the model designers were aware that features could be correlated with demographic groups in a way that would make them proxies. There's a huge problem with people trying to use umbrella usage to predict flooding. Some people are trying to develop a computer model that uses rainfall instead, but watchdog groups have raised concerns that rainfall may be used as a proxy for umbrella usage. (It seems rather strange…

They correctly note the existence of a tradeoff, but I don't find their statement of it very clear. Ideally, a model would be fair in the senses that: 1. In aggregate over any nationality, people face the same probability of a false positive. 2. Two people who are identical except for their nationality face the same probability of a false positive. In general, it's impossible to achieve both properties. If the output…

> Two people who are identical except for their nationality face the same probability of a false positive

It would be immoral to disadvantage one nationality over another. But we also cannot disadvantage one age group over another. Or one gender over another. Or one hair colour over another. Or one brand of car over another.

So if we update this statement:

> Two people who are identical except for any set of properties face the same probability of a false positive.

With that new constraint, I don't believe it is possible to construct a model which outperforms a data-less coin flip.

Re: We investigated Amsterdam's attempt to build a 'fair' fraud detection model

#43
post #5

The article talks a lot about fairness metrics but never mentions whether the system actually catches fraud. Without figures for true positives, recall, or financial recoveries, its effectiveness remains completely in the dark. In short: great for moral grandstanding in the comments section, but zero evidence that taxpayer money or investigative time was ever saved.

It also doesn't mention what numbers we are even talking about that given the expansive size of the Dutch government make this an at all useful thing.

Re: We investigated Amsterdam's attempt to build a 'fair' fraud detection model

#45

Congrats Amsterdam: they funded a worthy and feasible project; put appropriate ethical guardrails in place; iterated scientifically; then didn’t deploy when they couldn’t achieve a result that satisfied their guardrails. We need more of this in the world.

[flagged]

Re: We investigated Amsterdam's attempt to build a 'fair' fraud detection model

#46
post #36
post #25

Earlier quoted context omitted.

The goal is to avoid penalizing people for their skin color, or for gender/sex/ethnicity/whatever. If some group have higher rate of welfare fraud, the fair/unbiased system must keep false positives for that group at the same level as for general population. Ideally there should be no false positives at all, because they are costly for people, who were marked wrongly, but sadly real systems are not like that. So thes…

>The goal is to avoid penalizing people for their skin color [...] That's not correct. The goal is to identify and flag fraud cases. If one group has a higher likelihood to perform that, then this will show up in the data. The solution should not be to change the data but educate that group to change their behavior. Please note that I have neither mentioned any specific group and do not have a specific group in mind.…

In practice, investigations tend to find the results for which the investigation was started. At the beginning of the article, it was also suggested that such investigations in Amsterdam found no higher rate of actual fraud amongst the groups which were targeted more frequently via implicit bias by human reviewers.

In North America, we know that white people use hard drugs at a slightly higher rate than non-whites. However, the arrest and conviction rate of hard drug users is multiples higher for non-white people than whites. (I mention North America because similar data exist for both Canada and the USA, but the exact ratios and which groups are negatively impacted differ.)

Similarly, when it comes to accusations of welfare fraud, there is substantial bias in the investigations of non-whites and there are deep-seated racist stereotypes (thanks for that, Reagan) that don't hold up to scrutiny especially when the proportion of welfare recipients is slightly higher amongst whites than amongst non-whites[1].

So…saying that the goal is to avoid penalizing people for [innate characteristics] is more correct and a better use of time. The city of Amsterdam already knew that its fraud investigations were flawed.

[1] In the US based on 2022 data, https://www.census.gov/library/stories/2022/05/who-is-receiv... shows that excluding Medicaid/CHIP, the rate of welfare is higher for whites.

Re: We investigated Amsterdam's attempt to build a 'fair' fraud detection model

#47
post #36
post #25

Earlier quoted context omitted.

The goal is to avoid penalizing people for their skin color, or for gender/sex/ethnicity/whatever. If some group have higher rate of welfare fraud, the fair/unbiased system must keep false positives for that group at the same level as for general population. Ideally there should be no false positives at all, because they are costly for people, who were marked wrongly, but sadly real systems are not like that. So thes…

>The goal is to avoid penalizing people for their skin color [...] That's not correct. The goal is to identify and flag fraud cases. If one group has a higher likelihood to perform that, then this will show up in the data. The solution should not be to change the data but educate that group to change their behavior. Please note that I have neither mentioned any specific group and do not have a specific group in mind.…

> The solution should not be to change the data but educate that group to change their behavior.

1. This is easier to say than to do.

2. In reality what you see is a correlation. If you try to educate all 20 year old females to not become a connected to organized crime CEOs of construction companies, your efforts will be wasted with 99% of these people, because they are either not connected to organized crime or are not going to become CEOs. Moreover the very your efforts will lead to a discrimination of 20 years old females, if not due to public perception of them, then because you've just increased difficulties for them of becoming a CEO.

> The goal is to identify and flag fraud cases.

Not quite. The goal is to reduce the amount of fraud cases. To identify and flag is a method of achieving that goal. But policymakers has a lot of other goals, like avoiding discrimination or reducing rate of murders. By focusing on one goal policymakers might undermine other goals.

As a side (almost methaphisical) note: it is one of the reason, why techies are bad at social problems. Their math education taught them to ignore all irrelevant details, when dealing with a problem, but society is a big complex system where everything is connected, so in general you can't ignore anything, because everything is relevant. But education have the upper hand, so techies tend to throw away as much complexity as it is needed to make the problem solvable. They will never accept that they don't know how to solve a problem.

Re: We investigated Amsterdam's attempt to build a 'fair' fraud detection model

#48
post #47
post #36

Earlier quoted context omitted.

>The goal is to avoid penalizing people for their skin color [...] That's not correct. The goal is to identify and flag fraud cases. If one group has a higher likelihood to perform that, then this will show up in the data. The solution should not be to change the data but educate that group to change their behavior. Please note that I have neither mentioned any specific group and do not have a specific group in mind.…

> The solution should not be to change the data but educate that group to change their behavior. 1. This is easier to say than to do. 2. In reality what you see is a correlation. If you try to educate all 20 year old females to not become a connected to organized crime CEOs of construction companies, your efforts will be wasted with 99% of these people, because they are either not connected to organized crime or are…

"your efforts will lead to a discrimination of 20 years old females"

I'd think that this is an extremely far-fetched example that fails at basic logic. Just because a very specific scenario will be flagged does not mean that this scenario is generalized to all CEOs, all females, all 20 year olds.

Re: We investigated Amsterdam's attempt to build a 'fair' fraud detection model

#49
Why is there so much focus on "fair" even when reality isn't?

Not all misdeeds are equally likely to be detected. What matter is minimizing the false positives and false negatives. But it sounds like they don't even have a base truth to be comparing it against, making the whole thing an exercise in bureaucracy.

Re: We investigated Amsterdam's attempt to build a 'fair' fraud detection model

#50
post #45

Congrats Amsterdam: they funded a worthy and feasible project; put appropriate ethical guardrails in place; iterated scientifically; then didn’t deploy when they couldn’t achieve a result that satisfied their guardrails. We need more of this in the world.

[flagged]

> because I don't even need to look at the data to know that some groups are more likely to commit fraud.

That is by definition prejudice: bias without evidence. Perhaps they want to avoid that.

Post reply on HN