Live data from Hacker News

Statistical Mistakes and How to Avoid Them

cs.cornell.edu

41–50 of 80 posts

Re: Statistical Mistakes and How to Avoid Them

#41
post #2

This is the insight that made statistics "click" for me many years ago: a statistical test answers one central question: what are the odds that the results you observed could have arisen by chance? If those odds are low, then you are justified in concluding that the results probably did not arise by chance, and so there must be some other explanation (usually, but not always, the causal hypothesis you are advancing).…

Unfortunately no. Very much no, even though it's widely believed that that is a good definition/intuition (and used in many places).

It's the odds of having that results due to chance, if the null hypothesis is true[0]. That latter part might sound pedantic, but the whole point is that we don't know how likely the null hypothesis is. If I test wheather the sun has just died[1] and get a p-value of 0.01 it's still very likely that this result is due to change (surely more than 1%)! We need a prior probability (i.e. bayesian statistics) to calculate the probability that the result was due to chance, that is why that partial definition is incomplete and actually very misleading. This point is subtle, but very important to really understand p-values.

Another way to look at it is: if we knew the probability that the result was due to chance we could also just take 1-p and have to probability of there actually being some effect, a probability that hypothesis testing cannot give us.

There is one nice property that hypothesis testing does have (and why presumably it's so widely used): if the idea you are testing is wrong (which actually means "null hypothesis true") you will most likely (1-p) not find any positive results. This is good, this means that if the sun in fact did not die, and use 0.01 as your threshold, 99% of the experiments will conclude that there is no reason to believe the sun has died. So hypothesis testing does limit the number of false positive findings. The xkcd comic is a bit misleading it this regard, yes it does highlight the limitations of frequentist hypothesis testing, but the scenario depicted is a very unlikely one, in 99% of the cases there would have been a boring and reasonable "No, the sun hasn't died".

For an incredibly interesting article about the difficulty of concluding anything definitive from scientific results I highly recommend "The Control Group is out of Control" at slatestarcodex[2].

[0] To be even more pedantic you would have to add "equal or more extreme", and "under a given model", but "if the null hypothesis is true" is by far the most important piece often missing.

[1] https://xkcd.com/1132/

[2] http://slatestarcodex.com/2014/04/28/the-control-group-is-ou...

Re: Statistical Mistakes and How to Avoid Them

#42
post #30
post #29

Earlier quoted context omitted.

Someone ought to formulate the Godwin's law analogue for CLT. Hint 1: what happens when variance does not exist or is too large to be considered finite for any practical purposes. Hint 2: Levy processes, stable distributions

Too cryptic to be helpful, sorry..

Ah I see, my apologies, I assumed Googling those keywords would be sufficient to connect the dots.

The main thing that I wanted to convey is that the consequences of CLT does not come for free. It is not remotely as widely applicable as it is made out to be. CLT is also not so narrow that you need IID random variables as is often claimed. Those assumptions can be relaxed substantially. What gets in the way in obtaining a Gaussian distribution in the limit is the requirement that the original distribution(s) have a finite variance. The (averaging) process may still converge to some limiting distribution but it would not be a Gaussian one. Levy measures, more precisely stable measures are that class and Gaussian is but one member of that class, the only one that has finite variance. It is slowly getting acknowledged that many natural processes in fact do not have a finite variance. Fractals are one such process

Re: Statistical Mistakes and How to Avoid Them

#44
post #41
post #2

This is the insight that made statistics "click" for me many years ago: a statistical test answers one central question: what are the odds that the results you observed could have arisen by chance? If those odds are low, then you are justified in concluding that the results probably did not arise by chance, and so there must be some other explanation (usually, but not always, the causal hypothesis you are advancing).…

Unfortunately no. Very much no, even though it's widely believed that that is a good definition/intuition (and used in many places). It's the odds of having that results due to chance, if the null hypothesis is true [0]. That latter part might sound pedantic, but the whole point is that we don't know how likely the null hypothesis is. If I test wheather the sun has just died[1] and get a p-value of 0.01 it's still ve…

> It's the odds of having that results due to chance, if the null hypothesis is true[0].

Yes, that's right. I don't know why you think this is at odds with what I said. In fact, I clarified this myself a few hours ago in a sibling comment:

https://news.ycombinator.com/item?id=13026907

Re: Statistical Mistakes and How to Avoid Them

#45
post #24
post #22

Earlier quoted context omitted.

How about you guys are both right, sort of. There are both Bayesian and Frequentist approaches in statistics! They represent very different methods to statistics, but they are also quite similar. My apologies, I couldn't find one link that gave a good description of Bayesian vs. Frequentist. Here are a couple links to get started: https://xkcd.com/1132/ http://jakevdp.github.io/blog/2014/03/11/frequentism-and-bay...…

That cartoon is actually a pretty good illustration of why frequentists are wrong.

All models are wrong. That much is entailed by 'models'.

There is no Zuul.

Re: Statistical Mistakes and How to Avoid Them

#46
post #38
post #32

Earlier quoted context omitted.

> Do you mean probably wrong? Nope. > Both approaches have pros and cons What are the pros of the frequentist approach?

You can justify that you did everything "by the book" and you don't have to make up priors. Now. You might read this as me saying that Bayesian statistics is all made up, and that's not what I'm saying. I'm saying that if you go Bayesian all the way, and your result is even remotely controversial, someone could easily challenge you by saying "this result would have been different with different priors, and why won't…

In other words: the advantage of frequentism is that it gives you better plausible deniability when you get the wrong answer. Did I get that right?

Re: Statistical Mistakes and How to Avoid Them

#47
post #21

Earlier quoted context omitted.

> Statistics can tell you whether a model is consistent with the data. Yes, that's true, but it badly misses the point. The power of statistics is to tell you when a model (the null hypothesis) is (most likely) inconsistent with the data so that you can confidently rule it out. Any finite data set is consistent with an infinite number of models, so knowing that a model and the data are consistent tells you absolutely…

Except such a data set is also inconsistent with an infinite number of models, so ruling one out via rejecting a null also provides practically no information value and moves us no closer to understanding. /devils advocate In practical terms, we're not interested in true models, but useful ones, so the description of a model's consistency with observed data is often the more useful metric in practice than rejecting n…

> Except such a data set is also inconsistent with an infinite number of models, so ruling one out via rejecting a null also provides practically no information value and moves us no closer to understanding.

No, that's not true, because experiments are not done in a (figurative) vacuum. They are done in the context of an explanatory theory that has already gone through a vigorous filter and shown to be consistent with the all prior experimental data and has better explanatory power than all of its competitors. It is only when more than one theory survives this filter that an experiment is done, and the experiment is designed specifically to distinguish between the surviving theories.

So while it is true that an experiment allows you to eliminate an infinite number of theories, it's irrelevant, because by the time the experiment is done nearly all of those theories have already been eliminated anyway.

Re: Statistical Mistakes and How to Avoid Them

#48
post #6

Earlier quoted context omitted.

It's actually P(data|null-hypothesis).

Yup, and the null hypothesis is a reduced model. Different words, same thing.

I have no idea what you mean by "reduced" but they are not the same thing. The difference in their explanatory power. A (scientific) hypothesis explains things. A null hypothesis simply says, "This hypothesis is wrong" but doesn't say why. It's a crucial distinction.

Re: Statistical Mistakes and How to Avoid Them

#49
post #12
post #5

I don't like how the article tries to push statistics on the reader. If a CS paper compares a pair of averages, then that gives certain information. If statistics can add to that, and make the results a little more precise, then that is nice. But by no means is it absolutely necessary. And statistics will not give a conclusive result either. I think that authors should use statistics when they see fit, and when it do…

Needless to say, I disagree. It can be straight-up misleading to report means without including a more nuanced view of the distribution. You don't need to use a bunch of fancy statistics, but you do need to consider whether your results could have arisen by random chance. That's not a distraction; it's accurately reporting what you found. Here's one frightening example of spurious performance results in CS: https://w…

> It can be straight-up misleading

It is only misleading if the reader doesn't understand statistics. There is, imho, nothing wrong with putting all your focus on the subject matter, and skipping the statistics while being frank about it.

Also, if you need statistics to show that your method is better than other methods, then perhaps your method is not really that much better.

Re: Statistical Mistakes and How to Avoid Them

#50
post #44
post #41

Earlier quoted context omitted.

Unfortunately no. Very much no, even though it's widely believed that that is a good definition/intuition (and used in many places). It's the odds of having that results due to chance, if the null hypothesis is true [0]. That latter part might sound pedantic, but the whole point is that we don't know how likely the null hypothesis is. If I test wheather the sun has just died[1] and get a p-value of 0.01 it's still ve…

> It's the odds of having that results due to chance, if the null hypothesis is true[0]. Yes, that's right. I don't know why you think this is at odds with what I said. In fact, I clarified this myself a few hours ago in a sibling comment: https://news.ycombinator.com/item?id=13026907

"what are the odds that the results you observed could have arisen by chance?"

If you say it like this it will very easily be misinterpreted. Once your results are in there are two cases: (1) either the null hypothesis is true and you got those results due to chance, or (2) the null hypothesis is false and there was some actual effect outside of the null hypothesis that helped you get the results.

Due to this it is very easy to interpret you statement as referring to the probability of (1).

Two two following definitions of p-values sound similar but are not:

[Correct] The probability of getting the results by chance if the null hypothesis is true P(Results|H0)

[Wrong] The probability that you got the results by chance and thus the null hypothesis was actually true P(H0|Results)

I'm not saying you didn't get it, but somebody reading what you wrote can very easily be fooled. And there are a lot of dead wrong definitions on the web[0][1][2][3].

[0] https://www.americannursetoday.com/the-p-value-what-it-reall...

[1] https://practice.sph.umich.edu/micphp/epicentral/p_value.php

[2] http://natajournals.org/doi/full/10.4085/1062-6050-51.1.04

[3] http://www.cdc.gov/des/consumers/research/understanding_scie...

Post reply on HN