Live data from Hacker News

Attacking discrimination with smarter machine learning

research.google.com

71–80 of 201 posts

Re: Attacking discrimination with smarter machine learning

#71

Earlier quoted context omitted.

Interestingly the EU has banned insurance underwriting based on gender even with all the actuarial data backing it up. Age is still fair game though.

Which made the insurance for women, who are traditionally safer drivers because not taking unless risks, go up. Well facts can't be racist but they are at the same time.

But if you were a man who was an excellent driver, would you like it that you get to pay more just because you were born into a gender that was "more reckless"?

Plus, women want equal rights, so why not be equal in this, too?

I think for things like insurance, or healthcare taxes (the model adopted by most countries) it's fair to normalize the price for everyone. Otherwise, the system doesn't really work.

Re: Attacking discrimination with smarter machine learning

#72
post #49

Earlier quoted context omitted.

I think it's useful to figure out why there are discrepancies between two groups. For example, let's take blacks in the US. The data tells you that a black person is more likely to be a criminal than a white person. There are two possible reasons for this: (1) blacks are more prone to crime, or (2) blacks are more likely to live in circumstances that make them criminals. With access to only anecdotal data, I strongly…

> The data tells you that a black person is more likely to be a criminal than a white person. There are two possible reasons for this: I'd like to add a third possible reason for your consideration. Since "criminality" i.e. guilt of committing a crime is determined after a process engaging the law enforcement and justice systems, we have to examine whether there are inherent biases in those systems that result in ske…

This is definitely a point worth raising.

I'll defend the parent comment by observing that we can reproduce this effect with non-judicial metrics, like "murder rate in majority foo-race communities". Assuming the reporting rate for murders is very high across the board, this escapes bias in both policing and conviction rates, and sends us back to other explanations.

Even so, the policing/judicial question is really important.

There are some sociology results (I haven't seen a replication failure?) suggesting that not only race but stereotypical appearance within race influences sentence duration. There are all kinds of low-level-illegal practices whose effective criminality is shaped entirely by policing - think of any "side hustle", where selling loose cigarettes leads to lots of charges, but selling moonshine (the closest analogy I can think of) leads to relatively few.

More obviously, incidental discovery and stop-and-frisk policing produce wildly imbalanced charge rates. Walking down a city street with a pocketknife and pot is always illegal, but the odds of it being a crime depend heavily on who you are.

I'm not saying anything new, but it's a soapbox worth getting up on. It's a big deal in race, but it's also worth remembering the power of selective policing whenever we talk about criminalizing something widespread.

Re: Attacking discrimination with smarter machine learning

#73
post #49

Earlier quoted context omitted.

I think it's useful to figure out why there are discrepancies between two groups. For example, let's take blacks in the US. The data tells you that a black person is more likely to be a criminal than a white person. There are two possible reasons for this: (1) blacks are more prone to crime, or (2) blacks are more likely to live in circumstances that make them criminals. With access to only anecdotal data, I strongly…

> The data tells you that a black person is more likely to be a criminal than a white person. There are two possible reasons for this: I'd like to add a third possible reason for your consideration. Since "criminality" i.e. guilt of committing a crime is determined after a process engaging the law enforcement and justice systems, we have to examine whether there are inherent biases in those systems that result in ske…

Blacks are disproportionally policed, leading to higher likelyhoods of being accused of a crime. Blacks are convicted more often than white people arrested for the same crime. When black people are convicted of a crime, they are more likely to be sentenced to incarceration compared to whites convicted of the same crime. Blacks generally get harsher sentences when compared to whites who've commited the same crime. https://www.americanprogress.org/issues/race/news/2012/03/13... http://www.huffingtonpost.com/kim-farbota/black-crime-rates-...

Re: Attacking discrimination with smarter machine learning

#74
post #49
post #25

>[...] concept called equal opportunity. Here, the constraint is that of the people who can pay back a loan, the same fraction in each group should actually be granted a loan. This does not seem fair to me, because if this is applied then your race (group) would determine your credit score threshold which feels discriminatory to me. I feel that, by definition, it is not discriminatory only if none of your attributes…

I think it's useful to figure out why there are discrepancies between two groups. For example, let's take blacks in the US. The data tells you that a black person is more likely to be a criminal than a white person. There are two possible reasons for this: (1) blacks are more prone to crime, or (2) blacks are more likely to live in circumstances that make them criminals. With access to only anecdotal data, I strongly…

[deleted]

Re: Attacking discrimination with smarter machine learning

#75
post #25

>[...] concept called equal opportunity. Here, the constraint is that of the people who can pay back a loan, the same fraction in each group should actually be granted a loan. This does not seem fair to me, because if this is applied then your race (group) would determine your credit score threshold which feels discriminatory to me. I feel that, by definition, it is not discriminatory only if none of your attributes…

But how do you do that? Race is baked into a lot of the attributes your classifier is going to find informative. Is someone a good credit risk? To decide, you look at features like past payment history, available balance, zip code, etc.

Past payment history: if you're black, prior discriminatory behaviors may have limited your ability to open credit accounts, and thus you have less history to go on. Available balance has the same reasoning. Zip code correlates with race.

I'd hope very few people are including an "is_black" feature in their classifiers. If you eliminate anything that is informative towards race though, you're likely going to have a classifier that doesn't work very well.

The problem is that we have datasets that have arisen from a history that included significant racism, both overt and latent. There is no way to separate those effects from the data. You either get an "optimal" classifier that is racially biased in ways we don't want, or you get one that intentionally gives up some perceived performance in favor of fairness.

Re: Attacking discrimination with smarter machine learning

#76
post #50

You realize you're advocating totalitarianism, right?

No, I think he's more advocating for a world where we don't encode the literal status quo into computer models that make decisions for us going forward, which is something that nobody of any political stripe is likely to want.

What's wrong with encoding the status quo? Why do we have to encode progressive extremism into computer models?

Being progressive for the sake of "going forward" sounds like misguided idealism. What do we do when progressive computer models lead us down a harmful path, do we simply play the typical liberal blame game and point fingers everywhere else while digging our head into sand?

Re: Attacking discrimination with smarter machine learning

#77

Earlier quoted context omitted.

So (assuming males are more costly to insure) either (a) females pay more than they should and are effectively subsidizing males or (b) males pay less than they should and the insurance companies will go broke..

It's OK because we subsidize their Health Care: men and women pay the same for the same age, but women use way more than men.

Citation?

Re: Attacking discrimination with smarter machine learning

#78
post #66

I agree that this should be a possibility, while I think the odds that millions of law enforcement officers and criminal justice faculty would spontaneously conspire against a particular race, I also think it's more likely than some genetic predisposition toward crime (i.e., the first option).

Now, this is his first offence, but a young man is found high on PCB walking down the street hitting cars with a baseball bat. It takes five cops and a significant struggle to arrest him.

Having read that, what would you assume is race and economic background is?

After criminal proceedings he was sentenced to community service.

Now, what would you assume is race and economic background is?

PS: Bias is insidious and really hard to control for.

Re: Attacking discrimination with smarter machine learning

#79

Earlier quoted context omitted.

> The data tells you that a black person is more likely to be a criminal than a white person. There are two possible reasons for this: I'd like to add a third possible reason for your consideration. Since "criminality" i.e. guilt of committing a crime is determined after a process engaging the law enforcement and justice systems, we have to examine whether there are inherent biases in those systems that result in ske…

This is studiable and has been studied. A study regarding police killings, for example: Do White Police Officers Unfairly Target Black Suspects? https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2870189&... Using a unique data set we link the race of police officers who kill suspects with the race of those who are killed across the United States. We have data on a total of 2,699 fatal police killings for the years…

There are a lot of different questions here, though. This study only addresses one.

There are lots of ways for the system to be racially skewed/biased without being a product of personal bias - so many that I think the focus on "racist police" makes it hard to recognize a lot of easily-provable problems.

Stop-and-frisk is my go-to example of a system that produces bias regardless of the race or biases of the officers involved. In theory, it's an efficient use of limited police resources, it can be implemented race-blind, and it "only catches criminals". There's room to talk about harassment of non-criminals, but at least regarding the people who get arrested its an understandable idea.

In practice, criminality is a product of conviction. Stop-and-frisk mostly catches 'possession' crimes like personal-use drugs and illegal weapons (and since a 3-inch pocketknife is illegal in many cities, we shouldn't mistake this for violent intent). As a result, living in a stop-and-frisk area massively increases your odds of being charged with a low-grade crime - it's not as though carrying marijuana or a Leatherman is rare among un-policed groups. Even if you attempt a crude race-blind implementation, like policing based on neighborhood crime rate, you end up with a vicious cycle where crime rates are high because enforcement is high.

So I think we do a disservice when we limit our discussion and investigation to officer bias. Even when it's not present, it's still easy to build an unequal system.

Re: Attacking discrimination with smarter machine learning

#80
post #24

Earlier quoted context omitted.

The difference with machine learning is that the model isn't designed by humans, through an actuarial process we can keep our brains wrapped around. It's a black box. We are OK attributing "crash risk" to young male drivers, because we can observe both that they are as a cohort statistically likely to crash and also understand why that would be the case. On the other hand, we're not comfortable with the idea that a c…

I dont think most ML algos are as black box as you think, and if there were such an algo that regularly had a problem with finding "irrelevant correlations" then it wouldnt be a very good ML algo and people wouldn't use it.

You're conflating algorithms with data to some degree. The algorithm is responsible for finding the correlation, but whether that correlation is "irrelevant" depends entirely on the relationship between the training data and the full (presumably unknown) distribution over that set, as well as on fuzzy social interpretions (e.g., we may simply decide by fiat that race is an irrelevant feature for deciding mortgage approvals because we want to enforce that as true within the system).
Post reply on HN