Live data from Hacker News

Machine Bias

propublica.org

11–20 of 61 posts

Re: Machine Bias

#11
post #10

According to propublicas own analysis, the claim of bias cannot be shown to be statistically significant. https://www.propublica.org/article/how-we-analyzed-the-compa... This article is terrible data journalism and probably deliberately misleading. Step 1: write down conclusion. Step 2: do analysis. Step 3: if analysis doesn't support conclusion, write down a bunch of anecdotes. Really, here's her R script: https://g…

They analyzed what they could -- the outcomes of the algorithm (recommendation) and the accuracy of those recommendations. They picked out specific examples, but the analysis was over the whole data set. I think you missed these relevant parts from the article: > We obtained the risk scores assigned to more than 7,000 people arrested in Broward County, Florida, in 2013 and 2014 and checked to see how many were charge…

Go read the description of the statistical analysis or just view their R notebook:

https://github.com/propublica/compas-analysis/blob/master/Co...

Their own analysis shows that (p ~= 0) that high and medium risk factors are predictive. They also showed that the racial bias terms (race_factorAfrican-American:score_factorHigh, etc) are probably not predictive (p > 0.05).

Your quotes are not evidence of bias, though I see how they might confuse an innumerate reader. It's interesting how good a job this article is doing confusing the innumerate - it's almost as if it was written to mislead without technically lying.

For example, black defendants being pegged as being more likely to commit crimes can be caused by one of two things: bias or perhaps black defends actually are more likely to commit crimes. According to ProPublica's own analysis (see race_factorAfrican-American), the latter is actually the case. This is true with p = 4.52e-06 - see line [36].

Re: Machine Bias

#12
post #6

Weapons of Math Destruction http://boingboing.net/2016/01/06/weapons-of-math-destruction... It's easy to hide agenda behind an algorithm; especially when the details of the algorithm are not publicly visible.

It's far easier to hide an agenda behind verbiage and anecdotes. Go read the author's actual statistical analysis:

https://github.com/propublica/compas-analysis/blob/master/Co...

In the statistical analysis (unlike the verbiage) she is completely unable to hide the lack of bias and the accuracy of the algorithm, all of which are clearly on display in line [36]. In contrast, her verbiage somehow conveys the exact opposite impression.

Re: Machine Bias

#13

Earlier quoted context omitted.

As you note, the algorithm for judicial discretion is unknown. The algorithm for this software is fully known, just kept from the public.

The validity of the algorithm can be - and apparently has been - reliably tested and been found to be useful and mostly unbiased. This analysis has been performed by both the algorithm's creators and highly adversarial third parties, such as the author of this article. Both found that whatever bias there is is small, and cannot be distinguished from random chance. For example, the author of this very article has done…

Lets be clear -- if the null hypothesis in this case is true (that there is no bias), and all other assumptions made are true, there is a slightly greater than 5.7% chance of obtaining this result (or something even more skewed). That's a great bar for publication of SCIENCE. It's not a great bar for hiding behind a proprietary algorithm used in sentencing.

People talk about misuse of p-values, but this takes the cake.

Re: Machine Bias

#14

According to propublicas own analysis, the claim of bias cannot be shown to be statistically significant. https://www.propublica.org/article/how-we-analyzed-the-compa... This article is terrible data journalism and probably deliberately misleading. Step 1: write down conclusion. Step 2: do analysis. Step 3: if analysis doesn't support conclusion, write down a bunch of anecdotes. Really, here's her R script: https://g…

(From my above reply too, as it applies here also):

Lets be clear -- if the null hypothesis in this case is true (that there is no bias), and all other assumptions made are true, there is a slightly greater than 5.7% chance of obtaining this result (or something even more skewed). That's a great bar for publication of SCIENCE. It's not a great bar for hiding behind a proprietary algorithm used in sentencing. People talk about misuse of p-values, but this takes the cake.

Re: Machine Bias

#15

According to propublicas own analysis, the claim of bias cannot be shown to be statistically significant. https://www.propublica.org/article/how-we-analyzed-the-compa... This article is terrible data journalism and probably deliberately misleading. Step 1: write down conclusion. Step 2: do analysis. Step 3: if analysis doesn't support conclusion, write down a bunch of anecdotes. Really, here's her R script: https://g…

(From my above reply too, as it applies here also): Lets be clear -- if the null hypothesis in this case is true (that there is no bias), and all other assumptions made are true, there is a slightly greater than 5.7% chance of obtaining this result (or something even more skewed). That's a great bar for publication of SCIENCE. It's not a great bar for hiding behind a proprietary algorithm used in sentencing. People t…

If you want to criticize the details of her analysis, go ahead. I'm solidly in the Bayesian camp and I agree with you 100%. What I'd have done is computed posteriors on all these coefficients and then computed bayes factors/probability of bias.

I'm confused though; the mood affiliation of your post somehow suggests that her less than perfect choice of a statistical methodology somehow supports her claims. Could you explain that? Or am I simply misunderstanding what you are trying to say?

Also, lets suppose we just take her own analysis at face value, and don't view it through the p-value lens. The maximum likelihood estimate suggests that even if this effect is not random chance, it's not very big. I.e., the "score factor high" estimate is >8x larger than the "score factor high, race = black" estimate. Isn't this really good? Do you really think the human biases that this algorithm mitigates are lower than this?

Lastly, what specific analysis would convince you that this algorithm is predictive and non-biased (or more realistically, not very biased)?

Re: Machine Bias

#16

Earlier quoted context omitted.

The validity of the algorithm can be - and apparently has been - reliably tested and been found to be useful and mostly unbiased. This analysis has been performed by both the algorithm's creators and highly adversarial third parties, such as the author of this article. Both found that whatever bias there is is small, and cannot be distinguished from random chance. For example, the author of this very article has done…

Lets be clear -- if the null hypothesis in this case is true (that there is no bias), and all other assumptions made are true, there is a slightly greater than 5.7% chance of obtaining this result (or something even more skewed). That's a great bar for publication of SCIENCE. It's not a great bar for hiding behind a proprietary algorithm used in sentencing. People talk about misuse of p-values, but this takes the cak…

I replied over there: https://news.ycombinator.com/item?id=11756449

Re: Machine Bias

#17
post #6

Weapons of Math Destruction http://boingboing.net/2016/01/06/weapons-of-math-destruction... It's easy to hide agenda behind an algorithm; especially when the details of the algorithm are not publicly visible.

It's far easier to hide an agenda behind verbiage and anecdotes. Go read the author's actual statistical analysis: https://github.com/propublica/compas-analysis/blob/master/Co... In the statistical analysis (unlike the verbiage) she is completely unable to hide the lack of bias and the accuracy of the algorithm , all of which are clearly on display in line [36]. In contrast, her verbiage somehow conveys the exact opp…

The data analysis you link to is by Jeff Larson, while the primary author of the article is Julia Angwin.

Larson is still the second author so it is certainly a big question how he can present data showing no statistical correlation between race and score then have his name on an article saying the exact opposite that is clearly pushing an agenda. And as noted, one where the owners of the publication are also involved in a competing risk assessment product.

Re: Machine Bias

#18
post #6

Weapons of Math Destruction http://boingboing.net/2016/01/06/weapons-of-math-destruction... It's easy to hide agenda behind an algorithm; especially when the details of the algorithm are not publicly visible.

It's far easier to hide an agenda behind verbiage and anecdotes. Go read the author's actual statistical analysis: https://github.com/propublica/compas-analysis/blob/master/Co... In the statistical analysis (unlike the verbiage) she is completely unable to hide the lack of bias and the accuracy of the algorithm , all of which are clearly on display in line [36]. In contrast, her verbiage somehow conveys the exact opp…

Uh... it's all right there in your link, across several sections that analyze specific parts of the data.

> Black defendants are 45% more likely than white defendants to receive a higher score correcting for the seriousness of their crime, previous arrests, and future criminal behavior.

> Women are 19.4% more likely than men to get a higher score.

> Most surprisingly, people under 25 are 2.5 times as likely to get a higher score as middle aged defendants.

> The violent score overpredicts recidivism for black defendants by 77.3% compared to white defendants.

> Defendands under 25 are 7.4 times as likely to get a higher score as middle aged defendants.

> [U]nder COMPAS black defendants are 91% more likely to get a higher score and not go on to commit more crimes than white defendants after two year.

> COMPAS scores misclassify white reoffenders as low risk at 70.4% more often than black reoffenders.

> Black defendants are twice as likely to be false positives for a Higher violent score than white defendants.

> White defendants are 63% more likely to get a lower score and commit another crime than Black defendants.

Calling out one specific section that doesn't show bias doesn't magically exonerate the rest.

Re: Machine Bias

#19
post #17

Earlier quoted context omitted.

It's far easier to hide an agenda behind verbiage and anecdotes. Go read the author's actual statistical analysis: https://github.com/propublica/compas-analysis/blob/master/Co... In the statistical analysis (unlike the verbiage) she is completely unable to hide the lack of bias and the accuracy of the algorithm , all of which are clearly on display in line [36]. In contrast, her verbiage somehow conveys the exact opp…

The data analysis you link to is by Jeff Larson, while the primary author of the article is Julia Angwin. Larson is still the second author so it is certainly a big question how he can present data showing no statistical correlation between race and score then have his name on an article saying the exact opposite that is clearly pushing an agenda. And as noted, one where the owners of the publication are also involve…

It's not quite right that he shows no correlation between race and score. There is a strong correlation between race and score. This correlation is caused by the fact that blacks have a high recidivism rate (p = 4.52e-6).

What the analysis shows is that once you know the predicted score of the algorithm, using race doesn't give you extra information. If the scores were biased then you could correct them by using racial information to undo the bias.

For more detail on that last bit, read the "What if measurements are biased?" section of my blog post: https://www.chrisstucchio.com/blog/2016/alien_intelligences_...

(The details differ a bit - I describe linear regression rather than cox models. But the basic idea is the same.)

Re: Machine Bias

#20
post #18

Earlier quoted context omitted.

It's far easier to hide an agenda behind verbiage and anecdotes. Go read the author's actual statistical analysis: https://github.com/propublica/compas-analysis/blob/master/Co... In the statistical analysis (unlike the verbiage) she is completely unable to hide the lack of bias and the accuracy of the algorithm , all of which are clearly on display in line [36]. In contrast, her verbiage somehow conveys the exact opp…

Uh... it's all right there in your link, across several sections that analyze specific parts of the data. > Black defendants are 45% more likely than white defendants to receive a higher score correcting for the seriousness of their crime, previous arrests, and future criminal behavior. > Women are 19.4% more likely than men to get a higher score. > Most surprisingly, people under 25 are 2.5 times as likely to get a…

None of these things are evidence of bias.

The algorithm is biased if it's giving the wrong score due to race or redundantly encoded race. To show that the algorithm is biased, you need to show that (score, race) pairs are more predictive than (score, ) singletons.

Line [36] and [46] both attempt to address this question. The only one of these which is statistically significant is "race_factorOther:score_factorHigh" in line [46].

The other things you bring up are interesting, but do not show bias. At best they show disparate impact which isn't remotely the same thing.

Post reply on HN