Live data from Hacker News

If correlation doesn’t imply causation, then what does? (2012)

michaelnielsen.org

41–50 of 72 posts

Re: If correlation doesn’t imply causation, then what does? (2012)

#41
This is the fundamental reason why general AI might not be possible on a computer without a body. To infer causality, you must form a hypothesis then design an experiment to confirm/deny it. Passively observing the world can't disambiguate between complex correlations or causality (even with a fancy calculus). You need action to learn the intricacies of the world.

Think how the discovery of electricity led to electronics. There is nothing like electronics in the natural world, the only route to electronics is via an iteratively refined causal model of the universe.

Re: If correlation doesn’t imply causation, then what does? (2012)

#42

It is dangerous to assume causality from any data alone. (Data and statistics are over-rated nowadays). You need to do the harder work of discovering the proper mathematical model (equation) relating explicitly the dependent (caused) variables to the independent (causing) variables. In the absence of such a verified and proven model, you just can not take a shortcut of pulling causality out of statistics, like a rabb…

But why is it dangerous, on balance, to make assumptions of causality from data and statistics alone?

Animals, such as rats and ravens, face this problem all the time, and yet they can meaningfully effect the world in such manner that would imply causal understanding, and a sensitivity towards the difference between mere correlation or a correlation with causal potential.

Humans do the same as well, naive people who have never learned about experimental design, or have never learned the concept of correlation, also make useful judgments on the causal model behind ordinary problems and events.

How did these machines make actionable judgments on causality with nothing more than noisy inputs to their sensory systems? Through what technique did they discern the difference between mere correlation, and a correlation with exploitable causality?

Re: If correlation doesn’t imply causation, then what does? (2012)

#43

Good reading in addition to Pearl himself is via Cosma Shalizi. Ref list: http://vserver1.cscs.lsa.umich.edu/~crshalizi/notebooks/caus... Chapters: http://www.stat.cmu.edu/~cshalizi/uADA/13/lectures/ch22.pdf http://www.stat.cmu.edu/~cshalizi/uADA/13/lectures/ch23.pdf http://www.stat.cmu.edu/~cshalizi/uADA/13/lectures/ch24.pdf Incidentally, Shalizi is a great source for going back to the basics. His course at CMU "Adv…

I concur. His notebook is a wonder of great resources on many (many many) subjects, with often interesting comments. http://vserver1.cscs.lsa.umich.edu/~crshalizi/notabene/

Re: If correlation doesn’t imply causation, then what does? (2012)

#44
Sometime ago I tried to come up with the simplest possible explanation for Simpson's paradox. This was the result:

1) Imagine that most women with a certain disease survive, while most men die.

2) Imagine that most women with the disease take a certain medicine, while most men don't.

3) Imagine that the medicine has absolutely no effect. Women just happen to have better innate resistance to the disease, and also just happen to buy the medicine more because it's marketed to women.

Now if you do a statistical analysis without counting men and women separately, you will conclude that the medicine is very correlated with survival!

Note that there's no way to know in advance that you should slice the population along such-and-such variables, which can be a lot more subtle than just gender. Also note that the example works even if the medicine has a slight negative effect, i.e. you can reverse the direction of correlations by choosing to slice or not to slice.

I think such results make it clear that you can't easily trust conclusions from statistics. One minute you're thinking that cholesterol causes heart disease, and the next minute you're asking yourself, what if cholesterol is part of the body's response to heart disease? That's why we need randomized controlled studies, and theories of causality.

Re: If correlation doesn’t imply causation, then what does? (2012)

#45
post #21

> We can’t have X causing Y causing Z causing X! At least, not without a time machine. You should be able to have cyclic graphs, though. W could cause X which causes Y which causes Z which causes more X.

Perhaps you just need to include time. Does X_1 causing Y causing Z causing Z_2 solve your problem without requiring cyclic graphs?

Yes, but then the nodes in the graph become even more disconnected from the concretes in reality that they are supposed to represent.

I am actually highly skeptical of the entire notion of the article, but I haven't had time to finish the article yet. My skepticism comes from the fact that causal relationships can be explained by identifying and understanding the causal factors at play and drawing relevant conclusions. For example, smoking causes lung cancer because exposure to toxic chemicals causes genetic mutations. Existing mathematics (such as basic statistics, including correlation) can be used as additional empirical evidence to support such an explanation.

Re: If correlation doesn’t imply causation, then what does? (2012)

#46

It is dangerous to assume causality from any data alone. (Data and statistics are over-rated nowadays). You need to do the harder work of discovering the proper mathematical model (equation) relating explicitly the dependent (caused) variables to the independent (causing) variables. In the absence of such a verified and proven model, you just can not take a shortcut of pulling causality out of statistics, like a rabb…

But why is it dangerous, on balance, to make assumptions of causality from data and statistics alone? Animals, such as rats and ravens, face this problem all the time, and yet they can meaningfully effect the world in such manner that would imply causal understanding, and a sensitivity towards the difference between mere correlation or a correlation with causal potential. Humans do the same as well, naive people who…

It is probably the ability to make the abstract jump from the data and its correlations to generalised laws that marks the chief difference between people on one hand and animals and machines on the other. Correlation is only useful up to a point. It shows that there may be some relationship between the variables but it says nothing about its nature. Was X caused by Y or was Y caused by X or were X,Y both caused by some unknown Z, was it all just an accidental data sample, was it significant, with how much doubt? Does the assigned significance in fact rely on an unwarranted assumption of some underlying population distribution? Just too many questions and no answers. Besides, causality is problematic enough (see non-aristotelian philosophies and/or quantum mechanics) even without trying to demonstrate it with statistics.

Re: If correlation doesn’t imply causation, then what does? (2012)

#47

It is dangerous to assume causality from any data alone. (Data and statistics are over-rated nowadays). You need to do the harder work of discovering the proper mathematical model (equation) relating explicitly the dependent (caused) variables to the independent (causing) variables. In the absence of such a verified and proven model, you just can not take a shortcut of pulling causality out of statistics, like a rabb…

But why is it dangerous, on balance, to make assumptions of causality from data and statistics alone? Animals, such as rats and ravens, face this problem all the time, and yet they can meaningfully effect the world in such manner that would imply causal understanding, and a sensitivity towards the difference between mere correlation or a correlation with causal potential. Humans do the same as well, naive people who…

> naive people who have never learned about experimental design, or have never learned the concept of correlation, also make useful judgments on the causal model behind ordinary problems and events.

They also are often... racist. Or hold whatever other stereotypes to heart. Racism is just a good example of an extreme position to hold which is often due to assumptions of causality.

"Lots of minorities are in prison. There is a high correlation between being a minority and being in prison. Therefore being a minority leads to being a criminal."

This completely ignores external reasons why minorities might end up in prison more often than others. For instance, it could be that minorities have an equal amount of criminal activity as the general population, but are more likely to end up in prison because of it. Correlation does not imply causation.

I think the number of social issues that arise due to assumptions of causation is quite high, actually, and often leads to poor decision making in policy. That is why it is "dangerous."

Re: If correlation doesn’t imply causation, then what does? (2012)

#48

It is dangerous to assume causality from any data alone. (Data and statistics are over-rated nowadays). You need to do the harder work of discovering the proper mathematical model (equation) relating explicitly the dependent (caused) variables to the independent (causing) variables. In the absence of such a verified and proven model, you just can not take a shortcut of pulling causality out of statistics, like a rabb…

But why is it dangerous, on balance, to make assumptions of causality from data and statistics alone? Animals, such as rats and ravens, face this problem all the time, and yet they can meaningfully effect the world in such manner that would imply causal understanding, and a sensitivity towards the difference between mere correlation or a correlation with causal potential. Humans do the same as well, naive people who…

"Through what technique did they discern the difference between mere correlation, and a correlation with exploitable causality?"

Evolution - that is, assumptions that are accurate are favoured since they are more likely to lead to the animal surviving, vs embracing spurious correlations which are likely to get you killed.

Of course, this can break down if we try to apply the cognitive rules of thumb we've evolved to new domains outside our original evolutionary scope - our difficulties in thinking about statistics and probability are a great example of this. The math is trivial, but it just doesn't fit our brains very well without a great deal of cultural scaffolding.

Re: If correlation doesn’t imply causation, then what does? (2012)

#49
post #37
post #33

Earlier quoted context omitted.

Here is my understanding from a non-expert. The obviousness is suposed to come from if `a` causes `b`, `b` can not have caused `a` because that would need some form of time travel. `b` could cause `a2` though which is very similar to `a` but is separated in time from `a`.

Yeah that makes sense. But then he uses the model wrongly, since he puts "hidden factor" "smoking" "lung cancer" in the vertices, and not "hidden factor (t=0)", "smoking(t=1)" and "lung cancer(t=2)". Further if used like that, then a possible hidden factor could actually be "lung cancer(t=0)", since it could (conceivably) cause both "smoking(t=1)" and "lung cancer(t=2)".

> But then he uses the model wrongly, since he puts "hidden factor" "smoking" "lung cancer" in the vertices, and not "hidden factor (t=0)", "smoking(t=1)" and "lung cancer(t=2)".

The direction of the arrows shows causation with also implies time has passed though the amount of time is not specified in the model he is using, though it could be added. He has a section latter in the article about why he left it absent earlier and covers how it could be included.

> Further if used like that, then a possible hidden factor could actually be "lung cancer(t=0)", since it could (conceivably) cause both "smoking(t=1)" and "lung cancer(t=2)".

If lung cancer actually caused smoking then yes that would a simple solution. If that was the answer though the very next question would be what caused "lung cancer(t=0)". Which "smoking(t=-1)" if existed would be suspect, and another "hidden factor"(s) of course as well.

Re: If correlation doesn’t imply causation, then what does? (2012)

#50
post #33
post #31

>Obviously, it’d make no sense to have loops in the graph: I stopped reading here. Not saying models without loops cannot be useful though.

Here is my understanding from a non-expert. The obviousness is suposed to come from if `a` causes `b`, `b` can not have caused `a` because that would need some form of time travel. `b` could cause `a2` though which is very similar to `a` but is separated in time from `a`.

[deleted]
Post reply on HN