Live data from Hacker News

Simpson’s Paradox (2016)

forrestthewoods.com

61–70 of 84 posts

Re: Simpson’s Paradox (2016)

#61
This is one of my favorite paradoxes too. Here's why:

"... given the same table, one should sometimes follow the partitioned and sometimes the aggregated data, depending on the story behind the data, with each story dictating its own choice. Pearl considers this to be the real paradox behind Simpson's reversal." [0]

[0]https://en.wikipedia.org/wiki/Simpson%27s_paradox

Re: Simpson’s Paradox (2016)

#62
So, this is the data that the wikipedia page on Simpson's Paradox cites for the Berkeley study, and that the author of the article has quoted:

                     Men              Women
    Department Applied  Admitted Applied  Admitted
    A          [825]    62%      108      [82%]
    B          [560]    63%      25       [68%]
    C          325      [37%]    [593]    34%
    D          [417]    33%      375      [35%]
    E          191      [28%]    [393]    24%
    F          [373]    6%       341      [7%]

Above, I've bracketed in each pair of columns a) the sex with the most applicants and b) the sex with the most admissions, in a department. If that data is really the Berkeley data, then it's clear that the bias is against the sex with the most applicants, rather than either men or women.

I can propose a mechanism for this kind of (with some abuse of terminology) selection bias. A department accepts some applications, then realises they've admitted too many applicants of one sex and start rejecting applicants from the dominant sex in an attempt to redress the balance. They make a mess of it and end up biased too far in the opposite direction than they originally started.

Also note that in 4 out of 6 departments, more men applied than women, explaining why more departments appear biased against men (provided my observation holds).

However, I can't be sure whether this is actually the original data because it's nowhere to be found on my pdf copy of the study (Sex bias in graduate admission) which I believe I got from here: https://homepage.stat.uiowa.edu/~mbognar/1030/Bickel-Berkele.... If anyone knows where this data actually comes from, I'd welcome a pointer.

Re: Simpson’s Paradox (2016)

#63

So, this is the data that the wikipedia page on Simpson's Paradox cites for the Berkeley study, and that the author of the article has quoted: Men Women Department Applied Admitted Applied Admitted A [825] 62% 108 [82%] B [560] 63% 25 [68%] C 325 [37%] [593] 34% D [417] 33% 375 [35%] E 191 [28%] [393] 24% F [373] 6% 341 [7%] Above, I've bracketed in each pair of columns a) the sex with the most applicants and b) the…

You need to look at the figures. The differences that support your argument are minor and within the margin for error. You could similarly concluded that women are just smarter across the board.

Re: Simpson’s Paradox (2016)

#64

So, this is the data that the wikipedia page on Simpson's Paradox cites for the Berkeley study, and that the author of the article has quoted: Men Women Department Applied Admitted Applied Admitted A [825] 62% 108 [82%] B [560] 63% 25 [68%] C 325 [37%] [593] 34% D [417] 33% 375 [35%] E 191 [28%] [393] 24% F [373] 6% 341 [7%] Above, I've bracketed in each pair of columns a) the sex with the most applicants and b) the…

As a separate comment, which might be controversial, I would like to call bullshit on the entire claim of the Berkeley study in particular (and not about Simpson's Paradox in general). In the "Berkeley data" (if that's what it is), it's clear again that men applied to most departments in larger numbers than women. The Berkeley data claims that because more women were admitted on a per-department basis, more departments were biased against men.

Now, picture this. Alice and Bob share a pizza. Alice takes 7 pieces and Bob takes 3 (he's on an intermittent fasting diet so he only eats every other slice). Alice eats 4 of her slices, Bob eats 3 of his. At the end, Alice turns to Bob and says "boy, you're such a glutton! You scoffed down all of your slices, but I still have 3 left".

Is that a fair comparison? Well, no. Alice starts out with almost double the slices than Bob. Bob eats less than Alice, but he's accused of stuffing his face because he eats a larger proportion of his smaller share.

Same with the Berkeley data. If that is the Berkeley data.

Re: Simpson’s Paradox (2016)

#65

So, this is the data that the wikipedia page on Simpson's Paradox cites for the Berkeley study, and that the author of the article has quoted: Men Women Department Applied Admitted Applied Admitted A [825] 62% 108 [82%] B [560] 63% 25 [68%] C 325 [37%] [593] 34% D [417] 33% 375 [35%] E 191 [28%] [393] 24% F [373] 6% 341 [7%] Above, I've bracketed in each pair of columns a) the sex with the most applicants and b) the…

You need to look at the figures. The differences that support your argument are minor and within the margin for error. You could similarly concluded that women are just smarter across the board.

I'm sorry, I don't understand your comment. What difference is minor? What is the margin for error? And how would I conclude what you say?

Re: Simpson’s Paradox (2016)

#66
post #20

Iirc, you can guard against simpson's paradox by designing/collecting balanced data

I thought the same; at least in the kidney stone story, the data wasn't balanced: treatment A was assigned a lot more "harder cases". Either the trial wasn't randomized or the data set size wasn't big enough.

Unless you are God. You will never be able to even properly know what to factor in. Actually doing the experiment and analysis is exponentially harder. It's like saying well who cares about p=np, if you want to decrypt Aes without the key just make a super fast computer.

Re: Simpson’s Paradox (2016)

#67

Earlier quoted context omitted.

You need to look at the figures. The differences that support your argument are minor and within the margin for error. You could similarly concluded that women are just smarter across the board.

I'm sorry, I don't understand your comment. What difference is minor? What is the margin for error? And how would I conclude what you say?

Men are only favourites by 1-2%. That's within the margin if error. Women are favourites by say 10% plus. The comment treats them the same, and base their theory on a binary concept. It's just bad logic and may even be a version of the Simpson paradox.

Re: Simpson’s Paradox (2016)

#68
post #37

Earlier quoted context omitted.

I hate to take “both sides” but in the absence of confounding by indication, you can often use propensity scoring within robust models to decrease these impacts. Mind you, the problem with non random and undetected sampling bias is that it can be subtle. See for example https://www.nytimes.com/2018/08/06/upshot/employer-wellness-...

Propensity scoring is a method of applying statistical controls. How does it address the issue of controls compounding measurement error?

That’s the whole point of doubly robust models. However, in the event of confounding by indication or sampling misspecification, my experience is that nothing can save you.

I am a rather strong proponent of randomized trials for this exact reason. (They can also have sampling bias, but some degree of noise is inevitable)

Re: Simpson’s Paradox (2016)

#69
post #45

Observation #2: the paradox is essentially describing statistical gerrymandering. :)

Came here to say this. Simpson’s paradox is exactly how gerrymandering works. It’s all about how the data is grouped.

Re: Simpson’s Paradox (2016)

#70

So, this is the data that the wikipedia page on Simpson's Paradox cites for the Berkeley study, and that the author of the article has quoted: Men Women Department Applied Admitted Applied Admitted A [825] 62% 108 [82%] B [560] 63% 25 [68%] C 325 [37%] [593] 34% D [417] 33% 375 [35%] E 191 [28%] [393] 24% F [373] 6% 341 [7%] Above, I've bracketed in each pair of columns a) the sex with the most applicants and b) the…

As a separate comment, which might be controversial, I would like to call bullshit on the entire claim of the Berkeley study in particular (and not about Simpson's Paradox in general). In the "Berkeley data" (if that's what it is), it's clear again that men applied to most departments in larger numbers than women. The Berkeley data claims that because more women were admitted on a per-department basis, more departmen…

I'm not quite sure I follow your complaint, but I think I might be disagreeing with you. A key lesson of Simpson's Paradox is you can't read stories into data without having a causal model derived from outside the data.

I can comfortably invent stories that are not inconsistent with the data for a wide range of scenarios:

1) Only the most capable women are applying to Dept A due to discrimination, so the data is evidence of discrimination.

2) Dept A is discriminating towards women (self evident, 80% vs 60% admissions).

3) Dept A is completely non-discriminatory and the assessors are unaware of the gender of applicants; the differences are due to personal choices w.r.t. education and social networks turning out to be proxies for gender.

No study this sort of data can detect gender bias. It can be used as evidence in a broader study that comes up with a causal model for how the admissions process works; but there is no getting around interviews and field observations.

Post reply on HN