Live data from Hacker News

If correlation doesn’t imply causation, then what does? (2012)

michaelnielsen.org

31–40 of 72 posts

Re: If correlation doesn’t imply causation, then what does? (2012)

#32
post #21

> We can’t have X causing Y causing Z causing X! At least, not without a time machine. You should be able to have cyclic graphs, though. W could cause X which causes Y which causes Z which causes more X.

Perhaps you just need to include time. Does X_1 causing Y causing Z causing Z_2 solve your problem without requiring cyclic graphs?

Re: If correlation doesn’t imply causation, then what does? (2012)

#33
post #31

>Obviously, it’d make no sense to have loops in the graph: I stopped reading here. Not saying models without loops cannot be useful though.

Here is my understanding from a non-expert. The obviousness is suposed to come from if `a` causes `b`, `b` can not have caused `a` because that would need some form of time travel.

`b` could cause `a2` though which is very similar to `a` but is separated in time from `a`.

Re: If correlation doesn’t imply causation, then what does? (2012)

#35
It is dangerous to assume causality from any data alone. (Data and statistics are over-rated nowadays). You need to do the harder work of discovering the proper mathematical model (equation) relating explicitly the dependent (caused) variables to the independent (causing) variables. In the absence of such a verified and proven model, you just can not take a shortcut of pulling causality out of statistics, like a rabbit out of a hat. Incidentally, the "Simpson's paradox" to which so much attention is given here, is a trivial illustration of the fact that you can not meaningfully add percentages from differing amounts. Something every school kid ought to learn.

Re: If correlation doesn’t imply causation, then what does? (2012)

#36

It is dangerous to assume causality from any data alone. (Data and statistics are over-rated nowadays). You need to do the harder work of discovering the proper mathematical model (equation) relating explicitly the dependent (caused) variables to the independent (causing) variables. In the absence of such a verified and proven model, you just can not take a shortcut of pulling causality out of statistics, like a rabb…

Can you expand on this idea? To me it seems fine to draw causal inferences from data alone.

Re: If correlation doesn’t imply causation, then what does? (2012)

#37
post #33
post #31

>Obviously, it’d make no sense to have loops in the graph: I stopped reading here. Not saying models without loops cannot be useful though.

Here is my understanding from a non-expert. The obviousness is suposed to come from if `a` causes `b`, `b` can not have caused `a` because that would need some form of time travel. `b` could cause `a2` though which is very similar to `a` but is separated in time from `a`.

Yeah that makes sense. But then he uses the model wrongly, since he puts "hidden factor" "smoking" "lung cancer" in the vertices, and not "hidden factor (t=0)", "smoking(t=1)" and "lung cancer(t=2)". Further if used like that, then a possible hidden factor could actually be "lung cancer(t=0)", since it could (conceivably) cause both "smoking(t=1)" and "lung cancer(t=2)".

Re: If correlation doesn’t imply causation, then what does? (2012)

#38

Correlation + plausible based on your knowledge of the world implies causation (obviously to the appropriate degree). It's the flip-side of extraordinary claims require extraordinary evidence. Facebook driving Greek debt is implausible and two vaguely shaped curves aren't enough. A formula that predicts to many decimal places over a fair period, prospectively, would be really weird but hard to ignore. Spanish debt, g…

"Facebook driving Greek debt is implausible" Yes, but in reality you don't know that Plausibility is a subjective measure, and while I would say that, yes, it can be a hint, you cannot disregard something merely because it's implausible

> a hint, you cannot disregard something merely because it's implausible

Coming back to Bayesian statistics, the word for this is prior, and its not that you disregard evidence, its that you can quantify both your existing beliefs about reality and the change in your beliefs according to the evidence you see.

Firstly, our hint: we have a prior assumption of the probability that the popularity of Facebook is driving up Greek debt (call it P(Fg)). Then, we observe a correlation between these two things. For the sake of argument, I'm going to make this 0.001 (I'd probably estimate less).

Now, once we see this correlation, we now need to calculate two things: 1. The probability of observing that correlation (call that P(C)). Note that the more extraordinary the correlation, the less probable it is, and the smaller this term would be. In this case, the graph matches vaguely, I'm going to give it a probability of 0.1.

2. Given a world where there is a causation, what is the probability we'd see this correlation (Q: I'm not 100% on this part). Now, Greek debt could plausibly be driven by other things, which would mask the Facebook effect, so there's no guarantee there would be a correlation. This term is called P(C | Fg), and I have no idea what value to give it. Let's try 0.5.

What we want to know is: P(Fg | C), that is, the probability of a connection given we have observed a correlation.

Boom! P(Fg | C) = P(C | Fg) x P(Fg) / P(C)

So our posterior probability (after observing this correlation) changes from 0.001 to 0.001 x 0.5 / 0.1 = 0.005

Re: If correlation doesn’t imply causation, then what does? (2012)

#39

It is dangerous to assume causality from any data alone. (Data and statistics are over-rated nowadays). You need to do the harder work of discovering the proper mathematical model (equation) relating explicitly the dependent (caused) variables to the independent (causing) variables. In the absence of such a verified and proven model, you just can not take a shortcut of pulling causality out of statistics, like a rabb…

Can you expand on this idea? To me it seems fine to draw causal inferences from data alone.

It depends where the data comes from. The gold standard in science is a controlled, randomized experiment - but obviously that standard is not always attainable.

Re: If correlation doesn’t imply causation, then what does? (2012)

#40

It is dangerous to assume causality from any data alone. (Data and statistics are over-rated nowadays). You need to do the harder work of discovering the proper mathematical model (equation) relating explicitly the dependent (caused) variables to the independent (causing) variables. In the absence of such a verified and proven model, you just can not take a shortcut of pulling causality out of statistics, like a rabb…

Can you expand on this idea? To me it seems fine to draw causal inferences from data alone.

I think it helps a lot if you have a plausible explanation for the chain of causes. For example, the correlation of dead people in houses with shitty stoves and furnaces in those houses is interesting, but once you discover that said stoves and furnaces leak carbon monoxide, AND that carbon monoxide binds to hemoglobin with much greater affinity than oxygen, AND that the affected people don't receive any specific warning symptoms when that happens, you suddenly have a reason to stop wondering whether or not they were committing suicide at a heightened rate because of poverty-related depression.
Post reply on HN