Live data from Hacker News

Simpson’s Paradox (2016)

forrestthewoods.com

71–80 of 84 posts

Re: Simpson’s Paradox (2016)

#71

Earlier quoted context omitted.

I'm sorry, I don't understand your comment. What difference is minor? What is the margin for error? And how would I conclude what you say?

Men are only favourites by 1-2%. That's within the margin if error. Women are favourites by say 10% plus. The comment treats them the same, and base their theory on a binary concept. It's just bad logic and may even be a version of the Simpson paradox.

I still do not understand. How are men "favourites by 1%-2%" and women "by 10% plus"? Favourites, for what?

And how did you calculate the margin of error for this study?

Re: Simpson’s Paradox (2016)

#72
post #70

Earlier quoted context omitted.

As a separate comment, which might be controversial, I would like to call bullshit on the entire claim of the Berkeley study in particular (and not about Simpson's Paradox in general). In the "Berkeley data" (if that's what it is), it's clear again that men applied to most departments in larger numbers than women. The Berkeley data claims that because more women were admitted on a per-department basis, more departmen…

I'm not quite sure I follow your complaint, but I think I might be disagreeing with you. A key lesson of Simpson's Paradox is you can't read stories into data without having a causal model derived from outside the data. I can comfortably invent stories that are not inconsistent with the data for a wide range of scenarios: 1) Only the most capable women are applying to Dept A due to discrimination, so the data is evid…

I'm not challenging Simpson's paradox, only the conclusion quoted in respect with the data in the above table (I'm still not sure where it came from).

Re: Simpson’s Paradox (2016)

#74
post #47

Earlier quoted context omitted.

I'm unclear: what was the incorrect claim you're saying the author made?

The parent is not saying the author made an incorrect claim. They are saying that the parent did not continue their argument to arrive at a conclusion that someone else had, the conclusion that causal models are what tells you when you can combine datasets and when you can't.

> causal models are what tells you when you can combine datasets and when you can't.

but then the causal model is subjective right? What if there are two different causal models, and a priori cannot be known which is the "true" one?

Can the selection of the causal model be used to justify the dataset, in order to push a particular agenda?

Re: Simpson’s Paradox (2016)

#75

Earlier quoted context omitted.

I'm sorry, I don't understand your comment. What difference is minor? What is the margin for error? And how would I conclude what you say?

Men are only favourites by 1-2%. That's within the margin if error. Women are favourites by say 10% plus. The comment treats them the same, and base their theory on a binary concept. It's just bad logic and may even be a version of the Simpson paradox.

Women are the favorite by 10%+ only for a single department. This is a _different_ fallacy, now...

Re: Simpson’s Paradox (2016)

#76
post #74

Earlier quoted context omitted.

The parent is not saying the author made an incorrect claim. They are saying that the parent did not continue their argument to arrive at a conclusion that someone else had, the conclusion that causal models are what tells you when you can combine datasets and when you can't.

> causal models are what tells you when you can combine datasets and when you can't. but then the causal model is subjective right? What if there are two different causal models, and a priori cannot be known which is the "true" one? Can the selection of the causal model be used to justify the dataset, in order to push a particular agenda?

Your job when analysing data is simply to enumerate the possibilities and assign likelihoods to them if possible. If two models fit equally well, you're supposed to write them both down in the hope that someone will collect further data to distinguish between them.

If you're cutting holes in your report for political reasons, that's just not doing the job. That's what pundits are paid to do, not (ideally at least) scientists. Fraud is easy to commit, and the fact that it's possible is not that hard of a philosophical issue.

Re: Simpson’s Paradox (2016)

#77
post #74

Earlier quoted context omitted.

> causal models are what tells you when you can combine datasets and when you can't. but then the causal model is subjective right? What if there are two different causal models, and a priori cannot be known which is the "true" one? Can the selection of the causal model be used to justify the dataset, in order to push a particular agenda?

Your job when analysing data is simply to enumerate the possibilities and assign likelihoods to them if possible. If two models fit equally well, you're supposed to write them both down in the hope that someone will collect further data to distinguish between them. If you're cutting holes in your report for political reasons, that's just not doing the job. That's what pundits are paid to do, not (ideally at least) sc…

How do you tell that a paper containing conclusions to support an agenda is written with correct scientific rigor, rather than fraud? Using Simpson's paradox, one can obfuscate their biases by making the desired conclusion drop out of the data.

Re: Simpson’s Paradox (2016)

#78

Earlier quoted context omitted.

Men are only favourites by 1-2%. That's within the margin if error. Women are favourites by say 10% plus. The comment treats them the same, and base their theory on a binary concept. It's just bad logic and may even be a version of the Simpson paradox.

I still do not understand. How are men "favourites by 1%-2%" and women "by 10% plus"? Favourites, for what? And how did you calculate the margin of error for this study?

First each subject you compare the chance of admission. For men when they have a higher chance of admission, even in their most advantaged subject they have a higher chance of admission of 4%. Women on the other hand have a 20%. You can't say that they are equivalent in the least. In terms of error margins, a few percent is common, from experience. You could do a stats 95 confidence style calculation.

Re: Simpson’s Paradox (2016)

#79
post #61

This is one of my favorite paradoxes too. Here's why: "... given the same table, one should sometimes follow the partitioned and sometimes the aggregated data, depending on the story behind the data, with each story dictating its own choice. Pearl considers this to be the real paradox behind Simpson's reversal." [0] [0] https://en.wikipedia.org/wiki/Simpson%27s_paradox

Not really a paradox, but you will like https://en.wikipedia.org/wiki/Anscombe%27s_quartet

(if you've not encountered it before, which I suspect is unlikely!)

Re: Simpson’s Paradox (2016)

#80
I sometimes wonder why people expect there to be any fixed, categorical semantic relationship between any set of numbers and set of natural language statements.

Very rarely do the words or the numbers cover even a tiny amount of the possible interpretations.

Post reply on HN