Live data from Hacker News

A Way to Detect Bias

paulgraham.com

61–70 of 224 posts

Re: A Way to Detect Bias

#61

On a simple mathematical basis, this is false. Consider two groups of candidates for a scholarship, A and B. We want to select all candidates that have an 80% or better chance of graduation. Group A comes from a population where the chance of graduation is distributed uniformly from 0% to 100% and group B is from one where the chance is distributed uniformly from 10% to 90%, with the same average but less variation i…

Great insight. I had the great fortune to take a course with Gary Becker where we went into some of the mathematics of college admissions. He made this precise point -- the variances of the particular populations you are looking at matter a great detail. He managed to build some pretty convincing models which provided a compelling narrative for "biases" that we seemed to observe in the real world, all with simple cha…

I had the great fortune to take a course with Gary Becker

Lucky you. He was great. I never had the chance to take a class from him but it would have been worth a Chicago winter to have the chance.

Re: A Way to Detect Bias

#62
> A couple months ago, one VC firm (almost certainly unintentionally) published a study showing bias of this type. First Round Capital found that among its portfolio companies, startups with female founders outperformed those without by 63%.

Well, they also said in their study:

> And we are not claiming that our data is representative of the industry...or even statistically significant.

Also, the wording is "startups with a female founder", not exclusively female founders... I think this is a detail that shouldn't be ignored.

And, the study doesn't show how many companies out of the 300 had female founders! Maybe it was just 1! They also say "Solo Founders do Much Worse Than Teams", so this is an important detail if there are no solo female teams ever backed in their firm! etc etc, the list goes on. Not exactly strong evidence to support the point PG is making, that bias would be easy to detect.

Measuring performance purely in terms of "how much money I make" is one way of doing it, but not the only way. And it wont cover the majority of jobs on the planet (how do you measure performance of someone who stacks shelves in a supermarket?)

Re: A Way to Detect Bias

#63

The phenomena of "stereotype threat" complicates this conclusion however: https://en.wikipedia.org/wiki/Stereotype_threat When a member of a group is primed with a stereotype that their group underperforms at a task, they are more likely to underperform. So there could be a selection process biased against a group, and a selected member could be an above-average performer otherwise but, because of work environment, b…

This is only a bias if the stereotype effect is stronger on measurements than in real life. If stereotype threat affects test performance and real performance the same way, then it means that the stereotyped group is truly inferior. Do you have evidence that stereotype threat hurts test performance more than real performance? (Of course, in a hypothetical world which only eliminated the stereotype, the group would ce…

Minority groups will generally underperform when tested by majority groups (i.e. hispanic student white proctor) according to studies though I'm citing from memory so I may be incorrect. Also minority groups have access/insight/credibility with certain consumer groups and cultures majority groups do not and there's no reason to believe this advantage would be measured appropriately by examiners not fluent in that culture. If real-life performance means 'business success' and 'measurement performance' means 'being funded when pitching a startup to a white VCs' then I think that criterion is at its face satisfied.

Re: A Way to Detect Bias

#64
post #55

Earlier quoted context omitted.

So the PG estimator is clearly problematic. I agree that the yummfajitas (YM) estimator looks to be consistent. In this case though, we're dealing with (small) finite sample sizes, so we need to come up with some sort of test statistic. What would the YM test be here? It seems tricky since you are dealing with a conditional distribution based on left-censored data. I'm also not aware of any difference-of-minimums tes…

I don't know of something to refer to, but I don't think the statistics are too hard. The test statistic would be exactly min(sample1) and min(sample2). Suppose the cutoff sample is distributed according to f(x)H(x-C). Then the probability of the minima of a sample exceeding C+e by random chance, assuming the null hypothesis, is p = (1-\int_C^{C+e}f(x) dx)^N. So now you have a frequentist hypothesis test. If you make…

Does that assume both samples are identically distributed and the only difference is the cutoff? If it does, then couldn't we just continue to do a difference of means test and still be consistent? If it doesn't, how do you handle identifying the cutoff minima and the two different distributions in a frequentist way?

Re: A Way to Detect Bias

#65
post #36
post #10

Graham's statement about the possible bias of First Round is unfounded. This was not any sort of a real study like Graham thinks and First Round clearly notes that. When the returns are as skewed as they are in venture capital ( http://www.sethlevine.com/archives/2014/08/venture-outcomes-... ), a small sample size and a simple analysis won't do. First Round even excluded their investment in Uber because it would skew…

Even if it were a statistically appropriate sample size and female founders still out performed male founders it still wouldn't exlcude other likely explanations other than bias. What if the culture in venture capital is more willing to assist and mentor female founders leading to greater success? In academia there is a women are selected over equally qualified men 2:1 for tenure positions now [1]. It is not unreason…

> female founders still out performed male founders

That's not what the study says, it says groups that have at least 1 female founder, not female founders.

Re: A Way to Detect Bias

#66
post #64

Earlier quoted context omitted.

I don't know of something to refer to, but I don't think the statistics are too hard. The test statistic would be exactly min(sample1) and min(sample2). Suppose the cutoff sample is distributed according to f(x)H(x-C). Then the probability of the minima of a sample exceeding C+e by random chance, assuming the null hypothesis, is p = (1-\int_C^{C+e}f(x) dx)^N. So now you have a frequentist hypothesis test. If you make…

Does that assume both samples are identically distributed and the only difference is the cutoff? If it does, then couldn't we just continue to do a difference of means test and still be consistent? If it doesn't, how do you handle identifying the cutoff minima and the two different distributions in a frequentist way?

The only assumption I need is that P_{f,g}([C,C+d]) >= h(d) > 0 for some arbitrary monotonic function h(d). This comes directly from the p-value formula.

I.e., for any d, there is a finite probability of finding an A or a B in [C,C+d]. I don't actually care what the shapes of f or g are at all beyond this - as long as this probability exists and is bounded below (in whatever class of functions f and g might be drawn from), it's all fine.

Re: A Way to Detect Bias

#67
I don't understand his point about First Round Capital showing their female founders did better than companies without female founders. What does that show? How do we know that female founders aren't simply better? Or maybe women are scared of applying, so out of women, only the best apply? In that case, the mere idea that there is a bias can cause "pre-selection" bias.

I lack the mathematics to prove this, but it seems that on the face of it, pg is simply wrong. Or I'm misreading terribly.

Tangentially: Speaking of bias, why doesn't YC publish information on their companies' tech choices? PG racked up a lot of inferred cachet (positive) by stating that use of Lisp gave them a huge advantage. Now that YC has data, they should be able to show how choice of technology correlates to performance.

Re: A Way to Detect Bias

#68

Earlier quoted context omitted.

Even if Asians perform better for external reasons, the selection process should account for that before the selection is made, and the admitted class should be roughly equal performers, as a group. Unless Asians have a very lumpy shaped performance curve across the group

This presupposes that we should not accept unequal representation among groups. I don't believe that, and I don't think it's an implicit Western value. People should be allowed to flourish according to their natural advantages. The value is to try not to make an early judgment and exclude people based on assumptions about their group's capability, whether that exclusion is based on the group's perceived disadvantage…

You misunderstood what I wrote. I didn't claim that a high performing subgroup should be throttled, I claimed that the Asians that pass an unbiased bar would be roughly as successful as the non-Asians that pass the same bar. I

Re: A Way to Detect Bias

#69

Earlier quoted context omitted.

Alternately, you could conclude that instead of "mediocre" Asians being excluded by bias, Asians have an external advantage that makes them perform better. Maybe it's cultural, since most Asians are taught a very strong work ethic and heavy emphasis is placed on formal schooling, succeeding, and fitting in. Maybe Asians are physically better adapted to that type of work, with brains that retain information more easil…

That doesn't actually have a large effect. An example in numpy: In [14]: x = norm(0.0,1).rvs(100000) In [15]: mean(x[where(x > 2.0)]) Out[15]: 2.3774795090391301 In [16]: y = norm(0.5,1).rvs(100000) In [17]: mean(y[where(y > 2.0)]) Out[17]: 2.4372124830289557 I.e., a difference in the mean of 0.5 sigma corresponds to 0.06 in Graham's test statistic. Graham is a little bit off - a better place to look for bias is the…

At https://news.ycombinator.com/item?id=10483861 gizmo shares an intuitive anecdote that matches your math.

Re: A Way to Detect Bias

#70
post #46

Earlier quoted context omitted.

> People make decisions based on social, cultural, and physical expectations of them, and there's not anything wrong with that. I couldn't disagree more. If the culture is plain chauvinism ("women belong in the kitchen not the boardroom") then there's everything wrong with that. All oppression throughout history is essentially "just culture", but that justifies nothing. Your biological reductionism is completely at o…

I didn't say "just culture"; I said the confluence of social, cultural, and physical factors. Why does "culture" develop? Because people are naturally evil and black-hearted? These things don't happen in a vacuum, they develop organically because they are the best way to support human and tribal propagation and prosperity. Perhaps some things can and should change, but things that are constant across nearly all succe…

Before you take your theory too far, you need to explain why it's OK that your theory implies that black people in America were best suited to be slaves, up until the day they weren't.
Post reply on HN