Live data from Hacker News

Psychology Journal Bans Significance Testing

sciencebasedmedicine.org

31–40 of 88 posts

Re: Psychology Journal Bans Significance Testing

#31

There was a question yesterday about Evidence Based Medicine vs Science Based Medicine. The SBM criticism of EBM is the over-reliance on Randomized Controlled Trials that meet p=0.05, without looking at the prior probability that a treatment would help. For example, EMB would say that if you have an RCT that shows that a lucky rabbit's foot works, then you have reasonable evidence to put that into practice. The issue…

> The SBM criticism of EBM is the over-reliance on Randomized Controlled Trials that meet p=0.05, without looking at the prior probability that a treatment would help. The strength of significance testing is that it purposely doesn't try to tell you how likely something is to be true, only how likely the data you got was the result of chance assuming the treatment is no better than placebo. You're still taking the pr…

You're still taking the prior probability into account when trying to figure out the truth, you're just not putting a number on it.

This is a danger sign - you are doing the same things the Bayesians do, just informally, less explicitly, and probably incorrectly.

The fact is that to make a good decision, eventually you need to compute a single number. This is an elementary fact of topology:

https://www.chrisstucchio.com/blog/2014/topology_of_decision...

That number will be based on some unproveable assumptions. That's a fact of Godel's incompleteness theorem, if nothing else. So given this, why is it "dubious" to make those assumptions explicit and obvious?

Re: Psychology Journal Bans Significance Testing

#32
post #20

Earlier quoted context omitted.

Let me replay it to you and see if I understand it, because I don't know if I do: You assume a null hypothesis that (usually) represents the status quo of no influence between the theory and the data. You then collect data. The p value then describes the probability of that data aligning with / being as a result of the null hypothesis. In other words, a p Is that correct? I may have minced terms there because my stat…

No there's a 5% chance the null hypothesis is correct. the probability your theory is correct is unknowable.

This is false. The null hypothesis is either correct or false. The veracity of the null hypothesis does not change as long as the experiments are repeated the same.

Re: Psychology Journal Bans Significance Testing

#34
This is very interesting for me. I recently switched from AI to Human Computer Interaction which is more cross discipline, specifically with a strong influence from psychology. I'll gladly admit that I don't have the best understanding of NHST but I think I understand it well enough.

Interestingly enough there was a post on HN somewhat recently about Bayesian alternatives (BEST). The paper that was linked was: ftp://ftp.sunet.se/pub/lang/CRAN/web/packages/BEST/vignettes/BEST.pdf

And the recommended book I settled on was by the author of that paper (Doing Bayesian Data Analysis)

I feel like I'm "ahead of the curve" thanks to HN :D

Re: Psychology Journal Bans Significance Testing

#35
post #20

I spent about a hour explaining p-values to a fellow graduate-level researcher a few weeks ago. I pretty much just kept rephrasing the definition in slightly different ways until the person finally got it. In undergrad, hypothesis testing was more or less taught as "do this inscrutable calculation and if the result is 0.05 or less, you win". The point is, in my experience, a lot of people really don't get p-values, e…

Let me replay it to you and see if I understand it, because I don't know if I do: You assume a null hypothesis that (usually) represents the status quo of no influence between the theory and the data. You then collect data. The p value then describes the probability of that data aligning with / being as a result of the null hypothesis. In other words, a p Is that correct? I may have minced terms there because my stat…

Just to add to what the others say, rejecting the null hypothesis is also not evidence that your alternative hypothesis is correct.

Often the hypothesis testing framework is stated something like:

    H0: µ = 0 (null hypothesis)  
    Ha: µ ≠ 0 (alternative hypothesis)
When you reject H0, it means that you can be somewhat confident that there was was some kind of distortion in your data that moved the mean away from (in this case) 0.

You can create a number of theories that purport to explain this mechanistically, but you'll often need particular setups like a randomized controlled trial that can eliminate alternative explanations. When you've eliminated all the competing reasonable hypotheses, there can only be one. If you can use that hypothesis to make non-obvious predictions, that's further proof that it's right as well.

Hypothesis testing is there to tell you when to take an effect seriously, it doesn't tell you whether your explanation is right outside of very carefully constructed circumstances (i.e. where if you see a particular effect, only one theory can explain it).

Re: Psychology Journal Bans Significance Testing

#37
post #7

The article is certainly correct that p-values and confidence intervals (or confidence sets, in multi-dimensional contexts) are widely misunderstood, not just in psychology or other social sciences, but in the hard sciences as well. The problem is even worse when you look outside of academia at common practices in more applied settings. As suggested, a good approach is to take p-values not as conclusive or decisive,…

Is the cautious approach then to treat a p-value in the absence of priors on the same level as a p-value in presence of unfavorable priors? When someone tests positive for a cancer test, the priors are known (probability of cancer in the general population is usually very low, and the false positive rate of the test may be relatively high), and so usually that first test is merely indication that further tests are needed. So when you don't know the prior and you observe a low p-value on something, isn't that just "preliminary research" that needs to be further confirmed with other methods or at least the same test but using other data?

Re: Psychology Journal Bans Significance Testing

#38

Earlier quoted context omitted.

> The SBM criticism of EBM is the over-reliance on Randomized Controlled Trials that meet p=0.05, without looking at the prior probability that a treatment would help. The strength of significance testing is that it purposely doesn't try to tell you how likely something is to be true, only how likely the data you got was the result of chance assuming the treatment is no better than placebo. You're still taking the pr…

You're still taking the prior probability into account when trying to figure out the truth, you're just not putting a number on it. This is a danger sign - you are doing the same things the Bayesians do, just informally, less explicitly, and probably incorrectly. The fact is that to make a good decision, eventually you need to compute a single number. This is an elementary fact of topology: https://www.chrisstucchio.…

> The fact is that to make a good decision, eventually you need to compute a single number.

Your linked blog post states that if you make a good decision, then there is a process computing a single number which is equivalent to your process. This is not equivalent to what you claim. As a matter of fact, it's the same kind of confusion that exists around the p-value.

It's not the case that a process explicitly computing such a number automatically makes good decisions, which is what you seem to claim implicitly.

Also, Gödel has nothing whatsoever to to with this.

Re: Psychology Journal Bans Significance Testing

#39

Earlier quoted context omitted.

You're still taking the prior probability into account when trying to figure out the truth, you're just not putting a number on it. This is a danger sign - you are doing the same things the Bayesians do, just informally, less explicitly, and probably incorrectly. The fact is that to make a good decision, eventually you need to compute a single number. This is an elementary fact of topology: https://www.chrisstucchio.…

> The fact is that to make a good decision, eventually you need to compute a single number. Your linked blog post states that if you make a good decision, then there is a process computing a single number which is equivalent to your process. This is not equivalent to what you claim. As a matter of fact, it's the same kind of confusion that exists around the p-value. It's not the case that a process explicitly computi…

Eventually you need to compute a number which is either above or below your go/no go threshold. That's the number I'm referring to.

I don't claim you can't arrive at it by some perfect heuristic. I merely claim that you are better off being explicit about your assumptions and formalizing your reasoning. That just makes mistakes more obvious, makes your strong assumptions more clear, and makes it more likely that you will correctly update your beliefs rather than incorrectly discounting/overvaluing evidence.

You are right about godel, it's a separate theorem I'm referring to which says you need unproveable axioms. I misremembered, sorry, wrote that before my coffee.

Re: Psychology Journal Bans Significance Testing

#40

I spent about a hour explaining p-values to a fellow graduate-level researcher a few weeks ago. I pretty much just kept rephrasing the definition in slightly different ways until the person finally got it. In undergrad, hypothesis testing was more or less taught as "do this inscrutable calculation and if the result is 0.05 or less, you win". The point is, in my experience, a lot of people really don't get p-values, e…

I understand p-values, but I always have real problems understanding the thing of 95% confidence interval not meaning 95% probability of the true parameter being in the interval. I once grasped it, but then I forgot the reason. And now I look at this paragraph:

"the problem is that, for example, a 95% confidence interval does not indicate that the parameter of interest has a 95% probability of being within the interval. Rather, it means merely that if an infinite number of samples were taken and confidence intervals computed, 95% of the confidence intervals would capture the population parameter"

And I say, "What? If out of every 100 random samples, in 95 of them the parameter is in the interval, then surely the probability of the parameter being in the interval is 95% by definition?"

Where was the catch? I remember there was one (which is enough for practical purposes because it means I won't say in a paper that the probability of the parameter being in the interval is 5%) but I feel dumb for not being able to see at first glance something that is supposed to be basic statistics...

Post reply on HN