Live data from Hacker News

The Mathematics of Paul Graham's Bias Test

chrisstucchio.com

1–10 of 84 posts

Re: The Mathematics of Paul Graham's Bias Test

#2
Very nicely done Chris, but the basic problem with Paul’s analysis is not the mathematics (this can be fixed as you have shown), but the underlying data. Any data set you could get to measure bias in the start-up world is too small and messy to tell you anything useful. No matter how sophisticated your analysis, if the data is garbage then all you will end up with is garbage.

This does even consider the problem of data dredging which First Round Capital engaged in.

Re: The Mathematics of Paul Graham's Bias Test

#3

Very nicely done Chris, but the basic problem with Paul’s analysis is not the mathematics (this can be fixed as you have shown), but the underlying data. Any data set you could get to measure bias in the start-up world is too small and messy to tell you anything useful. No matter how sophisticated your analysis, if the data is garbage then all you will end up with is garbage. This does even consider the problem of da…

Maybe First Round's data set is too small or messy to get meaningful results, but the entire startup world has plenty of data to potentially pull some meaningful conclusions about bias.

I wonder if all YC companies would be enough data points to learn something useful. Or maybe grab a large swath of VC-funded startups including First Round's investments and many other top firms.

Most of the raw stats in Chris's post was above my head, but I'd love to see this applied to a larger data set of fundings.

Re: The Mathematics of Paul Graham's Bias Test

#4
Sorry, I still believe that a better approach is in

https://news.ycombinator.com/item?id=10484602

That post shows that what PG is doing is a first-cut effort at a statistical hypothesis test but with being vague on assumptions and without any information on false alarm rate.

In particular, in my post, get to compare sample averages without making a distribution assumption. Indeed make no distribution assumptions at all.

Yes, distributions exist, but that does not mean that we have to consider their details in all applications!

Come on guys, this is distribution-free statistical hypothesis testing, and we should be able to use that.

Re: The Mathematics of Paul Graham's Bias Test

#5
post #3

Very nicely done Chris, but the basic problem with Paul’s analysis is not the mathematics (this can be fixed as you have shown), but the underlying data. Any data set you could get to measure bias in the start-up world is too small and messy to tell you anything useful. No matter how sophisticated your analysis, if the data is garbage then all you will end up with is garbage. This does even consider the problem of da…

Maybe First Round's data set is too small or messy to get meaningful results, but the entire startup world has plenty of data to potentially pull some meaningful conclusions about bias. I wonder if all YC companies would be enough data points to learn something useful. Or maybe grab a large swath of VC-funded startups including First Round's investments and many other top firms. Most of the raw stats in Chris's post…

The tests here don't work when zero-plus-noise is the modal return, no matter how much data you throw at them.

Re: The Mathematics of Paul Graham's Bias Test

#6
Lots of math in here premised on shaky foundations:

>Group A comes from a population where the chance of graduation is distributed uniformly from 0% to 100% and group B is from one where the chance is distributed uniformly from 10% to 90%

>The mean of group B is not lower because of bias (which would be reflected near x=80), but because the very best members of group B are simply not as good as the very best members of group A.

Yes, if we can assume some a-priori knowledge about certain "groups" of people, then we can make a more "informed" decision. That's pretty much the definition of bias, isn't it? Paul Graham's point, as I understood it, was that those assumptions are often invalid. Therefore, bias could cause the market to under value someone or some company. Your counterpoint seems to be, "let's suppose those biases are legitimate."

Re: The Mathematics of Paul Graham's Bias Test

#7
> The idea is generally correct - bias in a decision process will be visible in post-decision distributions

I find what's wrong with the idea more fundamental, that it talks only about the 'selection process' but in fact bias that impacts success or failure can come at other points.

Re: The Mathematics of Paul Graham's Bias Test

#8

Lots of math in here premised on shaky foundations: >Group A comes from a population where the chance of graduation is distributed uniformly from 0% to 100% and group B is from one where the chance is distributed uniformly from 10% to 90% >The mean of group B is not lower because of bias (which would be reflected near x=80), but because the very best members of group B are simply not as good as the very best members…

Read it again. He's talking about the counterexample there. It's a hypothetical.

Re: The Mathematics of Paul Graham's Bias Test

#9
post #3

Very nicely done Chris, but the basic problem with Paul’s analysis is not the mathematics (this can be fixed as you have shown), but the underlying data. Any data set you could get to measure bias in the start-up world is too small and messy to tell you anything useful. No matter how sophisticated your analysis, if the data is garbage then all you will end up with is garbage. This does even consider the problem of da…

Maybe First Round's data set is too small or messy to get meaningful results, but the entire startup world has plenty of data to potentially pull some meaningful conclusions about bias. I wonder if all YC companies would be enough data points to learn something useful. Or maybe grab a large swath of VC-funded startups including First Round's investments and many other top firms. Most of the raw stats in Chris's post…

The basic problem is getting hold of good data. Most VC rightly consider their data in this area very valuable and they are not going to part with it easily. Not even Paul delved into YC’s data.

More fundamentally even if you could get enough data, the data is just too messy to analyse and draw any valid conclusions.

Re: The Mathematics of Paul Graham's Bias Test

#10
One other thing which both this and PG's original theory get wrong:

Their basic premise is wrong, if bias continues to exist after the selection event in question.

For example, if YC had (hypothetically) a real bias against black or women entrepreneurs, it is almost certain that future funding rounds, as well as all possible exit scenarios, would exhibit very much of the same bias.

In which case, the future "performance" of those candidates would be poor, and by PG's definition unbiased even though the only meaningful result is that YC is no more biased than subsequent performance evaluations.

Post reply on HN