Live data from Hacker News

Machine Bias

propublica.org

41–50 of 61 posts

Re: Machine Bias

#41
post #29

Earlier quoted context omitted.

"Oh sure, this algorithm is much more likely to have false positives on blacks, and much more likely to have false negatives on whites, and the results are that blacks are more likely to treated more harshly by the system. But it's not biased because of the definition of bias I'm using!" Orwell would have loved "disparate impact isn't bias" :)

Clearly statistics terminology is confusing you. The definition of bias is E[\hat{\theta} - \theta]. The definition of disparate impact is a predictor computing different means/quantiles for different protected classes. https://en.wikipedia.org/wiki/Bias_of_an_estimator https://en.wikipedia.org/wiki/Disparate_impact To understand this intuitively, here's a simple thought experiment. Consider Captain Hindsight, a pred…

Clearly the idea that words have meanings beyond statistics terminology is confusing you.

And no, I'm not calling standard mathematical terminology Orwellian. What I'm calling Orwellian is your describing a biased system as unbiased (by attempting to reframe the discussion around a specific statistical definition, chosen by you).

Re: Machine Bias

#42
post #38

Earlier quoted context omitted.

If you want to criticize the details of her analysis, go ahead. I'm solidly in the Bayesian camp and I agree with you 100%. What I'd have done is computed posteriors on all these coefficients and then computed bayes factors/probability of bias. I'm confused though; the mood affiliation of your post somehow suggests that her less than perfect choice of a statistical methodology somehow supports her claims. Could you e…

> maximum likelihood That may be grounds for a mistrial. Decisions about crimes are not judged by the "maximum likelihood". > what specific analysis would convince you that this algorithm is predictive and non-biased What is it going to take to convince you that the choice of model and which data to use as input is just as important as the analysis itself? > race_factor Depending on the situation, using race or other…

What is it going to take to convince you that the choice of model and which data to use as input is just as important as the analysis itself?

I'm already convinced of this. Are you trying to imply that the cox model is wrong or something? If so, why not just make that argument explicitly?

Of course, if the Cox model is wrong, why do you believe the algorithm is biased? Isn't that reason to disregard the entire ProPublica article (which is all based on the Cox model)?

Depending on the situation, using race or other protected classes is illegal.

Did you even read the article? "Northpointe’s core product is a set of scores derived from 137 questions that are either answered by defendants or pulled from criminal records. Race is not one of the questions. "

/sigh/

You can emote all you like. Reality does not change.

I must admit, the emotion on display here confuses me. Much like you I oppose racial bias. The R script provides evidence that very little racial bias is present in this system. Why does this inspire such negative emotion? It's almost as if you care more about looking anti-racist than you care about having racism's effects be reduced.

Re: Machine Bias

#43
post #30
post #27

Earlier quoted context omitted.

> but do not show bias. At best they show disparate impact I have no interest in playing but-what-does-the-exact-dictionary-definition-say semantics games.

The funny thing is that the dictionary definition supports your point, not yummyfajitas. bias: prejudice in favor of or against one thing, person, or group compared with another, usually in a way considered to be unfair.

If you are going to argue a single dictionary definition, you should immediately stop holding a mouse when using a computer. Rodents don't like being held for long times, be connected electrically to a machine, nor do they like being rubbed on a pad.

Re: Machine Bias

#44
post #10

Earlier quoted context omitted.

They analyzed what they could -- the outcomes of the algorithm (recommendation) and the accuracy of those recommendations. They picked out specific examples, but the analysis was over the whole data set. I think you missed these relevant parts from the article: > We obtained the risk scores assigned to more than 7,000 people arrested in Broward County, Florida, in 2013 and 2014 and checked to see how many were charge…

Go read the description of the statistical analysis or just view their R notebook: https://github.com/propublica/compas-analysis/blob/master/Co... Their own analysis shows that (p ~= 0) that high and medium risk factors are predictive. They also showed that the racial bias terms (race_factorAfrican-American:score_factorHigh, etc) are probably not predictive (p > 0.05). Your quotes are not evidence of bias, though I s…

[deleted]

Re: Machine Bias

#45
post #36

Earlier quoted context omitted.

Your thought experiment here is incorrect, given that the analysis compares COMPAS results to actual recidivism rates and shows over- and under-prediction in comparison to them.

The thought experiment is a mathematical proof that the two concepts are causally unrelated, nothing more. I really suggest you brush up on your basic math - you seem to not be following along. Your claims about the emirical means of recidivism rates do not prove what you think they prove. Different races might be misclassified at different rates for a variety of reasons - e.g., one race might be affected more by som…

> Could you clearly lay out the statistical argument that you believe implies that E[\hat{\theta} - \theta] > 0?

Again: I don't care about your domain-specific definition of "bias". I care about whether the end result of this secret algorithm is unequal and inaccurate treatment of different demographic groups.

Re: Machine Bias

#46
post #45

Earlier quoted context omitted.

The thought experiment is a mathematical proof that the two concepts are causally unrelated, nothing more. I really suggest you brush up on your basic math - you seem to not be following along. Your claims about the emirical means of recidivism rates do not prove what you think they prove. Different races might be misclassified at different rates for a variety of reasons - e.g., one race might be affected more by som…

> Could you clearly lay out the statistical argument that you believe implies that E[\hat{\theta} - \theta] > 0? Again: I don't care about your domain-specific definition of "bias". I care about whether the end result of this secret algorithm is unequal and inaccurate treatment of different demographic groups.

inaccurate treatment of different demographic groups.

This is exactly what the standard mathematical definition of bias (restricted to a given group) addresses. The authors of this article ran exactly that analysis - see lines [36] and [46].

I know that you are trying to retreat from statistics, since the stats don't support your mood affiliation, but don't retreat to "accuracy". Retreat to something vague and undefined instead. It'll work better.

As for "unequal", I don't know what you mean. Do you consider disparate impact to be "unequal"? If so, then I'm sorry to tell you that reality is imposing an unfortunate choice on you: equal or accurate, you can't have both. (According to ProPublica this algorithm chooses accurate.)

Criticizing an algorithm for revealing this unfortunate fact is like blaming telescopes for Saturn having rings.

Re: Machine Bias

#47

One of the most mind boggling sentences in that article was: "On Sunday, Northpointe gave ProPublica the basics of its future-crime formula — which includes factors such as education levels, and whether a defendant has a job. It did not share the specific calculations, which it said are proprietary." How on earth can you lock people up based on secret information? That is Kafka meets Minority Report.

I'm leaning towards a Constitutional Amendment against automated law. The wording escapes me (and I'm unqualified anyhow) but the gist would be that only humans can judge humans, no machinery can be allowed to do it.

https://en.wikipedia.org/wiki/Butlerian_Jihad

Re: Machine Bias

#48
post #25
post #3

This is totally fucked. Morally wrong, deeply unethical, and probably illegal – if you're adding punishment without having that additional punishment based on new evidence, isn't that like being treated guilty without proof? Obviously I'm not a lawyer, but how could anyone, let alone the whole huge set of people that led to these policies, think that applying group statistics to individuals to determine the severity…

Punishment is always considered somewhat separately from the determination of guilt. The judge would already try to account for things like this when determining your sentence. They just do it in a deeply ad hoc and personal manner, where they just take a stab at it, try to account for things like how sorry you seem to be, apply guidelines, and come up with a number. This means that you might ultimately be punished f…

All this does is systematize those biases so that they can't be challenged like a judge with a record of bias can. The statistics that they choose to record create bias in and of themselves - by using race in the algorithm, you are building in the possibility that race influences criminality. If you built in favorite foods, some foods would end up resulting in higher sentences than others, just as if you built in phases of the moon when the crime was committed or the astrological sign of the victim.

Where there was absolutely no effect, one out of every twenty combinations of all other variables would show significance in combination with the current value of that particular variable in the likelihood of future crime.

Furthermore, the algorithm would simply extend existing biases in arrest and sentencing, because it simply can't account for crimes that are uncaught and unpunished. Groups that are stopped, searched, arrested, and convicted at greater rates would without fail be sentenced to more time. Just another benefit of being white in America.

You end up using the fact that some groups are punished more often to justify punishing them more harshly.

Even worse, I bet that the fact that it thinks that women are at a higher risk for recidivism means that somewhere within the algorithm it's using the fact that women in general are less criminal than men to decide that women who do commit crime are more exceptional (within women), and therefore more deviant. It's disgusting. If you can't legally discriminate against a person on particular grounds, you certainly can't feed those grounds into an algorithm to let it discriminate for you while you shrug and feign innocence.

The algorithm is the innocent one - it's just attempting to reflect the system as it is. It's like an algorithm you would write to predict the winners of horse races, or the sports book. And just like one of those algorithms, if you stuff it with garbage (the kind of garbage that makes it wrong 77% of the time), it will result in garbage. If you use the results for something not external to the system, bad variables will feed back into themselves and make the results progressively worse - what's the effect of a longer sentence on recidivism? How does profitable is the arbitrage on your sports book algorithm if people use the results to bet, and the distribution of bets shift the odds?

Re: Machine Bias

#49
post #45

Earlier quoted context omitted.

> Could you clearly lay out the statistical argument that you believe implies that E[\hat{\theta} - \theta] > 0? Again: I don't care about your domain-specific definition of "bias". I care about whether the end result of this secret algorithm is unequal and inaccurate treatment of different demographic groups.

inaccurate treatment of different demographic groups. This is exactly what the standard mathematical definition of bias (restricted to a given group) addresses. The authors of this article ran exactly that analysis - see lines [36] and [46]. I know that you are trying to retreat from statistics, since the stats don't support your mood affiliation, but don't retreat to "accuracy". Retreat to something vague and undefi…

> equal or accurate, you can't have both

By "equal and accurate", I mean a result that corresponds to actual recidivism rates for as many demographic groups as possible (males, females, different races, different ages, combinations of the above, etc). The analysis shows that it is substantially likely (certainly well-beyond the oft-vaunted "reasonable doubt" standard, nevermind the specifics of p-values being slightly above 0.5) that the algorithm being used doesn't provide such a result.

Re: Machine Bias

#50
post #43
post #30

Earlier quoted context omitted.

The funny thing is that the dictionary definition supports your point, not yummyfajitas. bias: prejudice in favor of or against one thing, person, or group compared with another, usually in a way considered to be unfair.

If you are going to argue a single dictionary definition, you should immediately stop holding a mouse when using a computer. Rodents don't like being held for long times, be connected electrically to a machine, nor do they like being rubbed on a pad.

I think the obvious answer here is to genetically engineer a "computer mouse" that can mentally work complex algorithms, as supplied via a pet training clicker.

Then, you'll be able to click your computer mouse to solve problems.

Post reply on HN