Live data from Hacker News

We investigated Amsterdam's attempt to build a 'fair' fraud detection model

lighthousereports.com

61–70 of 87 posts

Re: We investigated Amsterdam's attempt to build a 'fair' fraud detection model

#61
What nobody seems to talk about is that their resulting models are basically garbage. If you look at the last provided confusion matrix, their model is right in about 2/3 of cases when it makes a positive prediction. The actual positives are about 60%. So, any improvement is marginal at best and a far cry from ~90% accuracy you would expect from a model in such a high-stakes scenario. They could have thrown a half of cases out at random and had about the same reduction in case load without introducing any bias into the process.

Re: We investigated Amsterdam's attempt to build a 'fair' fraud detection model

#62

Earlier quoted context omitted.

I think you took too much of a jump, considering all properties the same, as if the only way to make the system fair is to make it entirely blind to the applicant. We tend to distinguish between ascribed and achieved characteristics. It is considered to be unethical to discriminate upon things a person has no control over, such as their nationality, gender, age or natural hair color. However, things like a car brand…

> It is considered to be unethical to discriminate upon things a person has no control over, such as their nationality, gender, age or natural hair color. Nationality and natural hair color I understand, but age and gender? A lot of behaviors are not evenly distributed. Riots after a football match? You're unlikely to find a lot of elderly women (and men, but especially women) involved. Someone is fattening a child?…

> A lot of behaviors are not evenly distributed.

That’s true. But the idea is that feeding it to a system as an input could be considered unethical, as one cannot control their age. Even though there’s a valid correlation.

> If you assume perfect free will, sure. But do you?

I’m not. If this matters, I’m actually currently persuaded that free will doesn’t exist. Which doesn’t change that if one buys a car, its make is typically all their decision. Whenever such decision is coming from them having a free will or entirely determined by antecedent causes doesn’t really matter for purposes of fraud detection (or maybe I fail to see how it does).

I mean, we don’t need to care why people do things (at all, in general) - it matters for how we should act upon detection, but not for detecting itself. And, as I understand it, we know we don’t want to cause unfair pressure on groups defined by factors they cannot change. Because when we did that it consistently contributed to various undesirable consequences. E.g. discrimination and stereotypes against women or men, or prejudice against younger or elder people didn’t do us any well.

Re: We investigated Amsterdam's attempt to build a 'fair' fraud detection model

#63

What nobody seems to talk about is that their resulting models are basically garbage. If you look at the last provided confusion matrix, their model is right in about 2/3 of cases when it makes a positive prediction. The actual positives are about 60%. So, any improvement is marginal at best and a far cry from ~90% accuracy you would expect from a model in such a high-stakes scenario. They could have thrown a half of…

> What nobody seems to talk about is that their resulting models are basically garbage.

The post does talk about it when it briefly mentions that the goal of building the model (to decrease the number of cases investigated while increasing the rate of finding fraud) wasn't achieved. They don't say any more than that because that's not the point they are making.

Anyway, the project was shelved after a pilot. So your point is entirely false.

Re: We investigated Amsterdam's attempt to build a 'fair' fraud detection model

#64

What nobody seems to talk about is that their resulting models are basically garbage. If you look at the last provided confusion matrix, their model is right in about 2/3 of cases when it makes a positive prediction. The actual positives are about 60%. So, any improvement is marginal at best and a far cry from ~90% accuracy you would expect from a model in such a high-stakes scenario. They could have thrown a half of…

You can't tell a project will fail until you undertake it.

Amsterdam didn't deploy their models when they found their outcome is not satisfactory. I find it a perfectly fine result.

Re: We investigated Amsterdam's attempt to build a 'fair' fraud detection model

#65

In my view, we need to move the goalposts. Fraud detection models will never be fair. Their job is to find fraud. They will never be perfect, and the mistaken cases will cause a perfectly honest citizen to be disadvantaged in some way. It does not matter if that group is predominantly 'people with skin colour X' or 'people born on a Tuesday'. What matters is that the disadvantage those people face is so small as to b…

Some groups will be more disadvantaged than others by being investigated. For example for welfare, I expect fraudsters to have more money to support themselves or less people to support (unless the criteria for welfare is something unexpected). So I'd say that there also needs to be more protections than just providing money.

Nevertheless the idea of giving money is still good imo, because it also incentivizes the fraud detection becoming more efficient, since mistakes now cost more. Unfortunately I have a feeling people might game that to get more money by triggering false investigations.

Re: We investigated Amsterdam's attempt to build a 'fair' fraud detection model

#66

What nobody seems to talk about is that their resulting models are basically garbage. If you look at the last provided confusion matrix, their model is right in about 2/3 of cases when it makes a positive prediction. The actual positives are about 60%. So, any improvement is marginal at best and a far cry from ~90% accuracy you would expect from a model in such a high-stakes scenario. They could have thrown a half of…

> What nobody seems to talk about is that their resulting models are basically garbage. The post does talk about it when it briefly mentions that the goal of building the model (to decrease the number of cases investigated while increasing the rate of finding fraud) wasn't achieved. They don't say any more than that because that's not the point they are making. Anyway, the project was shelved after a pilot. So your p…

Good catch about the project being shelved. It is buried pretty deep in the document to the point of making it misleading:

> In late November 2023, the city announced that it would shelve the pilot.

I would agree that implications regarding the use of those models do not hold, but not the ones about their quality.

Re: We investigated Amsterdam's attempt to build a 'fair' fraud detection model

#67
> None of these features explicitly referred to an applicant’s gender or racial background, as well as other demographic characteristics protected by anti-discrimination law. But the model designers were aware that features could be correlated with demographic groups in a way that would make them proxies.

What's the problem with this? It isn't racism, it's literally just Bayes' Law.

Re: We investigated Amsterdam's attempt to build a 'fair' fraud detection model

#68

Earlier quoted context omitted.

> Two people who are identical except for their nationality face the same probability of a false positive It would be immoral to disadvantage one nationality over another. But we also cannot disadvantage one age group over another. Or one gender over another. Or one hair colour over another. Or one brand of car over another. So if we update this statement: > Two people who are identical except for any set of properti…

I think you took too much of a jump, considering all properties the same, as if the only way to make the system fair is to make it entirely blind to the applicant. We tend to distinguish between ascribed and achieved characteristics. It is considered to be unethical to discriminate upon things a person has no control over, such as their nationality, gender, age or natural hair color. However, things like a car brand…

Could we look at what kind of achieved characteristics exists that do not act as a proxy for an ascribed characteristics, because I have a really hard time to find those. Culture and values are highly intertwined with behavior, and the bigger the impact the behavior has on a person life, it seems that the stronger the proxy behavior is going to be.

To take a few examples, looking at employment characteristics will have a strong relationship with gender, generally creating greater false positives for women. Similarly, academic success will have greater false positives for men. Where a person choose to live will proxy heavily towards social economic factors, which in turn has gender as a major factor.

Welfare fraud in itself also has differences between men and women. The sums tend to be higher for men. Women in turn dominate the users of the welfare system. Women and men also tend to receive welfare at different time in their life. It possible even that car brand has a correlation with gender which then would act as a proxy.

In terms of defining fairness, I do find it interesting that the Analogue Process gave men a beneficial advantage, while both the initial and the reweighed model are the opposite and give women an even bigger beneficial advantage. The change in bias against men created by using the detection algorithms is actually about the same size as the change in bias against non-dutch nationality between initial model and the reweighed one.

Re: We investigated Amsterdam's attempt to build a 'fair' fraud detection model

#69

> None of these features explicitly referred to an applicant’s gender or racial background, as well as other demographic characteristics protected by anti-discrimination law. But the model designers were aware that features could be correlated with demographic groups in a way that would make them proxies. What's the problem with this? It isn't racism, it's literally just Bayes' Law.

Let's say you are making a model to judge job applicants. You are aware that the training data is biased in favor of men, so you remove all explicit mentions of gender from their CVs and cover letters.

Upon evaluation, your model seems to accept everyone who mentions a "fraternity" and reject anyone who mentions a "sorority". Swapping out the words turns a strong reject into a strong accept, and vice versa.

But you removed any explicit mention of gender, so surely your model couldn't possibly be showing an anti-women bias, right?

Re: We investigated Amsterdam's attempt to build a 'fair' fraud detection model

#70
post #69

> None of these features explicitly referred to an applicant’s gender or racial background, as well as other demographic characteristics protected by anti-discrimination law. But the model designers were aware that features could be correlated with demographic groups in a way that would make them proxies. What's the problem with this? It isn't racism, it's literally just Bayes' Law.

Let's say you are making a model to judge job applicants. You are aware that the training data is biased in favor of men, so you remove all explicit mentions of gender from their CVs and cover letters. Upon evaluation, your model seems to accept everyone who mentions a "fraternity" and reject anyone who mentions a "sorority". Swapping out the words turns a strong reject into a strong accept, and vice versa. But you r…

I've never had any implication of my gender other than my name in any CV over the past decade.

Who are these people who make a career history doc include gender-implicating data? And if there are such CVs, they should be stripped of such data before processing.

The fraternity example is such a specific 1 in a 1000 case.

Post reply on HN