Live data from Hacker News

A Way to Detect Bias

paulgraham.com

81–90 of 224 posts

Re: A Way to Detect Bias

#81

People, most of whom clearly are not that good at math, are being really harsh on Paul Graham. Graham is mostly right, but slightly incorrect. In particular, suppose group A has the distribution f(x) and B has the distribution g(x). If f(x) and g(x) are shaped significantly differently past the cutoff , then mean(H(x-C)f(x)) and mean(H(x-c)g(x)) might not agree even though there is no bias by construction. (Here H(x)…

Isn't this not as meaningful as the mean, because the minimum doesnt really tell you about true capability? I can also easily construct an institution that satisfies your parameters, the minimums of two groups match, as in they have a token member at C for every group, but the characteristic performance of all members is different?

This in my eyes amounts to a simplification that is neat mathematically, but removes most of the useful information. I admit though that if the minimums did not match very well, that is an obvious sign of bias.

Edit: I suppose what I'm essentially saying is, what if minimums do match, then we cannot rule out bias.

Re: A Way to Detect Bias

#83
post #73

On a simple mathematical basis, this is false. Consider two groups of candidates for a scholarship, A and B. We want to select all candidates that have an 80% or better chance of graduation. Group A comes from a population where the chance of graduation is distributed uniformly from 0% to 100% and group B is from one where the chance is distributed uniformly from 10% to 90%, with the same average but less variation i…

The article defines bias as follows: > Want to know if the selection process was biased against some type of applicant? Check whether they outperform the others. This is not just a heuristic for detecting bias. It's what bias means. Under that definition, you have been biased against A. [edit: on reflection I see this as a weakness of his definition. I missed that your selection process does in fact select the best c…

Yes, but that's not the common usage of the word. Or how most people understand it. With that usage you could say ivy league schools are NOT biased against Asians since Asian graduates aren't more successful than non-Asian ones, except nobody does.

Re: A Way to Detect Bias

#84
post #79

Earlier quoted context omitted.

The only assumption I need is that P_{f,g}([C,C+d]) >= h(d) > 0 for some arbitrary monotonic function h(d). This comes directly from the p-value formula. I.e., for any d, there is a finite probability of finding an A or a B in [C,C+d]. I don't actually care what the shapes of f or g are at all beyond this - as long as this probability exists and is bounded below (in whatever class of functions f and g might be drawn…

Sorry, I'm confused here. A p-value makes an implicit assumption that your null hypothesis is a known N(0,1). That may be throwing me off a bit. I get the point of you want to look at the likelihood function which is just one minus the CDF in the given interval. I'm just not clear on how you can get around f and g being arbitrarily parameterized functions of a given class. Are you assuming we know the class and somet…

A null hypothesis is just a specific thing you are trying to disprove. In this case, it's simply that the min of both distributions is identical.

I am assuming we know exactly one thing about the class the measures f and g come from: for every function in that class, \int_C^{C+d} f(x) dx >= h(d) for some monotonic function h(d).

The p-value is then computed in terms of h(d), since p >= h(d)^N.

Re: A Way to Detect Bias

#85

On a simple mathematical basis, this is false. Consider two groups of candidates for a scholarship, A and B. We want to select all candidates that have an 80% or better chance of graduation. Group A comes from a population where the chance of graduation is distributed uniformly from 0% to 100% and group B is from one where the chance is distributed uniformly from 10% to 90%, with the same average but less variation i…

The problem here is language and what our actual objectives are.

When people complain about bias, they are not really talking about mathematical bias, but about something else: Their idea of fairness. They are talking about discrimination. And when we are discussing that, we can't really think about whether rules are applied fairly or not, but whether the rules produce the outcomes that we want.

Let's go for a ludicrous example: We'll accept all applicants whose IQ is higher than their weight in pounds. We'll be explicitly discriminating against heavy people, but at the same time, we have pretty clear implicit biases against men, and ethnic groups who tend to be taller. We might as well have said that we prefer children and Japanese women. There's no need for mathematical bias: The bias comes from the rule selection.

So, in your example, if our actual objective is to graduate an even amount of people from groups A and B, we have to, explicitly, make it easier for group B to get the scholarship. And many times organizations have objectives like that.

As a more real example, let's consider a police department. If the objective is to have a racial makeup that represents the community, and different races have different drop-out rates, the candidate selection will prefer one kind over the other, precisely to counter the drop-out differential.

So when regular people, and not mathematicians, discuss bias, the mathematical definition is unimportant. The one important thing is our stated objectives.

Re: A Way to Detect Bias

#86
post #74

I don't get it :-/ Why is there a bias? Even if the VCs are totally unbiased, why couldn't the startups with women outperformed the others? It could happen for a variety of reasons. Just hypothetically speaking, maybe startups-with-women have different networking connections or insight that male-only-startups don't have?

Then, if I understand PG's argument correctly, the VCs should invest in even more startups-with-women, and lower their threshold on investing them, at least until they perform no better than startups-without-women.

Re: A Way to Detect Bias

#87
post #74

I don't get it :-/ Why is there a bias? Even if the VCs are totally unbiased, why couldn't the startups with women outperformed the others? It could happen for a variety of reasons. Just hypothetically speaking, maybe startups-with-women have different networking connections or insight that male-only-startups don't have?

If that were the case, VCs should accept more companies-with-women, which would change the TOTAL mix of accepted candidates so that the women no longer overperformed.

The issue, interestingly, is that there are a lot of women (or biased-against group) between the median performance and 160% performance who are being rejected.

In other words, only the best women can get accepted, which means that above-median (but not superstars) are getting rejected, while many more above-median men are getting accepted.

This has the ring of truth to me (as pg says, it's the definition of bias).

Re: A Way to Detect Bias

#88
post #74

I don't get it :-/ Why is there a bias? Even if the VCs are totally unbiased, why couldn't the startups with women outperformed the others? It could happen for a variety of reasons. Just hypothetically speaking, maybe startups-with-women have different networking connections or insight that male-only-startups don't have?

Then, if I understand PG's argument correctly, the VCs should invest in even more startups-with-women, and lower their threshold on investing them, at least until they perform no better than startups-without-women.

Indeed!

Re: A Way to Detect Bias

#89
PG:

Assumption: There is no fundamental difference between a female and a male founder for achieving start-up success (average rates and variance/distribution of rates is the same)

Observation: VC funded start-ups with female founders are (on average) 60% more successful than start-ups with male founders

Hypothesis: VC funding is biased against female founders. The ones that do receive funding are better vetted, less risky, and have higher individual qualities.

Experiment: Start funding more female founders.

If we then observe: The numbers start to even out, then there is no fundamental difference. VC funding bias may have been the cause of the difference in success rate.

If we then observe: The numbers stay the same, then there is a fundamental difference and our assumption is flawed.

Rational choice: Start funding more female founders. This either removes a bias (levels the playing field), or increases your profit (funding more potentially successful founders).

PG should of course not use an hypothesis to prove an assumption (experiment/probing is needed for verification). But also: The possibility of an uneven distribution should not invalidate such an experiment (or PG's line of reasoning), it will merely bring it to light (the numbers would stay the same, thus we have shown that the difference is fundamental and not caused by a sampling bias).

Re: A Way to Detect Bias

#90

I don't understand his point about First Round Capital showing their female founders did better than companies without female founders. What does that show? How do we know that female founders aren't simply better? Or maybe women are scared of applying, so out of women, only the best apply? In that case, the mere idea that there is a bias can cause "pre-selection" bias. I lack the mathematics to prove this, but it se…

The argument is that First Round Capital must have implictly made it harder for female founders to get funding, since the ones who do perform better. The rational course of action for First Round Capital would be to lower their threshold on female founders (or, conversely, raise the threshold on male founders) until they perform no better or no worse than male founders.
Post reply on HN