Live data from Hacker News

It’s not just p=0.048 vs. p=0.052

statmodeling.stat.columbia.edu

61–70 of 92 posts

Re: It’s not just p=0.048 vs. p=0.052

#61
post #33

Related, about p-values: > Here's the problem in a nutshell: If you run 1000 experiments over the course of your career, and you get a significant effect (p > […] However, this is a statement about what happens when the null hypothesis is actually true. In real research, we don't know whether the null hypothesis is actually true. If we knew that, we wouldn't need any statistics! In real research, we have a p value, a…

For about a year or so now Ive been wanting to make a game about science. It'd basically be a research and discovery simulator, and there would be a free-play mode. Some of the knobs would be # of required replication, required p-value and how good people are at generating hypothesis. I think it'd be eye opening.

Please do. We're getting closer and closer to the tipping point (although it still might be far away) where people go "Oooohhhh... Well, shit.". Where we all realise. Your game would be one of the many catalysers for this.

This is my opinion of course, it's not like I have any scientific basis for this.

Re: It’s not just p=0.048 vs. p=0.052

#62
post #33

Related, about p-values: > Here's the problem in a nutshell: If you run 1000 experiments over the course of your career, and you get a significant effect (p > […] However, this is a statement about what happens when the null hypothesis is actually true. In real research, we don't know whether the null hypothesis is actually true. If we knew that, we wouldn't need any statistics! In real research, we have a p value, a…

For about a year or so now Ive been wanting to make a game about science. It'd basically be a research and discovery simulator, and there would be a free-play mode. Some of the knobs would be # of required replication, required p-value and how good people are at generating hypothesis. I think it'd be eye opening.

I hope you do! I think simulations can be a powerful way to teach/learn and I would like to see this one.

I put a reminder to check back with you about this in a few months. Is Keybase your preferred contact method? I know it in name only, but I imagine I can figure it out.

Re: It’s not just p=0.048 vs. p=0.052

#63

> Also, to get technical for a moment, the p-value is not the “probability of happening by chance.” Is it not? According to Wikipedia, it's "[...] the probability that, when the null hypothesis is true, the statistical summary [...] would be equal to, or more extreme than, the actual observed results." This sounds pretty much like "probability of happening by chance".

Close, it's the probability of it happening by chance given that the null hypothesis is true. Unless H0 is objectively and absolutely true, it is not the same as "probability of it happening by chance."

Re: It’s not just p=0.048 vs. p=0.052

#64

> Also, to get technical for a moment, the p-value is not the “probability of happening by chance.” Is it not? According to Wikipedia, it's "[...] the probability that, when the null hypothesis is true, the statistical summary [...] would be equal to, or more extreme than, the actual observed results." This sounds pretty much like "probability of happening by chance".

It sounds pretty much the same, but is distinctly different. That's where much of the confusion in the popular press comes from.

The difference is that, as highlighted in your quote, there is some null hypothesis that is assumed when discussing p-values.

For example: what is the probability of drawing x>2 when the underlying distribution is assumed to be a standard normal distribution N(0,1)?

The probability is small in this case, and could provide evidence to reject the null hypothesis (i.e. the distribution is not standard normal). It doesn't tell you about the probability of drawing x>2, it only gives evidence to reject (or not) the null hypothesis.

The wiki has more elaborate explanation. And probably better examples than mine.

Re: It’s not just p=0.048 vs. p=0.052

#66

For somebody who has never had formal training in statistics and similar discussions, what is a good/book to start groking these concepts?

If you want a college textbook, there are literally hundreds with titles similar to "introduction to probability and statistics" and you should choose the cheapest second hand book you can find (or that they have in your library). Content is mostly the same, of course the writing style will be different and some may be more appealing to you personally. Search engines are your friend for narrowing down your list.

For a popsci work, you could check out "how to lie with statistics", a classic.

https://en.m.wikipedia.org/wiki/How_to_Lie_with_Statistics

Re: It’s not just p=0.048 vs. p=0.052

#67

Put a Number on It! did a piece a while ago going through some psychology pieces that were part of a replication effort. They found that half failed to replicate but that people in a betting market could often tell which ones were going to replicate or not. The author also did a blind test himself and was also able to guess which ones would replicate. He laid out several rules of thumb, most significantly to the arti…

That is a pretty interesting conclusion, do you know whether his work has been replicated to confirm the outcomes?

Yep, it's confirmed with p = 0.049

Re: It’s not just p=0.048 vs. p=0.052

#68
post #33

Related, about p-values: > Here's the problem in a nutshell: If you run 1000 experiments over the course of your career, and you get a significant effect (p > […] However, this is a statement about what happens when the null hypothesis is actually true. In real research, we don't know whether the null hypothesis is actually true. If we knew that, we wouldn't need any statistics! In real research, we have a p value, a…

>If you run 1000 experiments over the course of your career, and you get a significant effect (p I'm all for criticism of p-values, but when I read a lot of critiques, I get to this point and simply stop reading.

No statistics text book that I've read assigns the magic value of p=0.05 and labels it as significant. All the ones I've read tell you to pick a p-value appropriate to your experiment. Yes, I get it that many social scientists don't have much of a clue and use 0.05 as some special threshold, but let's direct the criticism to the guilty parties, instead of blaming a statistical methodology.

I mean, we all know people who misuse the mean ("the average number of breasts a person has is 1") and ignore the shape of the distribution and the standard deviation. Yet we don't say "Let's stop using the mean!"

Re: It’s not just p=0.048 vs. p=0.052

#69
post #33

Related, about p-values: > Here's the problem in a nutshell: If you run 1000 experiments over the course of your career, and you get a significant effect (p > […] However, this is a statement about what happens when the null hypothesis is actually true. In real research, we don't know whether the null hypothesis is actually true. If we knew that, we wouldn't need any statistics! In real research, we have a p value, a…

This blog post is extremely misleading -- it appears to confuse the "null hypothesis" with an actual scientific hypothesis that is being tested. For example: "However, there are many cases where I am testing bold, risky hypotheses—that is, hypotheses that are unlikely to be true."

Those may be the hypotheses being tested, but they are not NULL hypotheses. NHST is not about testing the NULL hypothesis, it is about testing a non-Null hypothesis (the null hypothesis should always be incredibly boring and expected). There may be problems, but this blog post does not describe one.

Re: It’s not just p=0.048 vs. p=0.052

#70

Put a Number on It! did a piece a while ago going through some psychology pieces that were part of a replication effort. They found that half failed to replicate but that people in a betting market could often tell which ones were going to replicate or not. The author also did a blind test himself and was also able to guess which ones would replicate. He laid out several rules of thumb, most significantly to the arti…

This is true in more than just psych. A strategy for systems papers is to select the system that is second on the graphs. The author's results are not trustworthy, but the second place system was usable and ran reasonably well in the hands of somebody other than the original author.
Post reply on HN