Live data from Hacker News

Psychology Journal Bans Significance Testing

sciencebasedmedicine.org

21–30 of 88 posts

Re: Psychology Journal Bans Significance Testing

#22

>The type of analysis being banned is often called a frequentist analysis I find that there is a trend of associating "bad statistics" with "Frequentists Statistics" which isn't really fair. If you found a statistician trained only in Frequentist methods and asked their opinion on experiment design in psychological research they would likely be just as appalled as any Bayesian. I'm a big fan of Bayesian methods, but…

[deleted]

Re: Psychology Journal Bans Significance Testing

#24

I spent about a hour explaining p-values to a fellow graduate-level researcher a few weeks ago. I pretty much just kept rephrasing the definition in slightly different ways until the person finally got it. In undergrad, hypothesis testing was more or less taught as "do this inscrutable calculation and if the result is 0.05 or less, you win". The point is, in my experience, a lot of people really don't get p-values, e…

I agree. I'm in the pool of people who don't really understand p-values even after taking 2 statistics courses. The main issue I've noticed is that compared to many other calculations, you just have no idea if your result is correct. You can make wrong assumptions, wrong calculation, apply wrong methods, but in the end you get a number... and that's it. It may be a wrong number and you'll never know. This is complete…

It's actually entirely possible to do some "checking of your answer" for p-values, as well. As you mentioned, for practical math, you can often validate your answer against the initial assumptions. This is true for statistical testing too, as it typically relies on many theoretical assumptions. So what you can do in practice is propose a different set of assumptions, perform the hypothesis test in a manner that follows those new assumptions, and see if you obtain a similar result. Typically, for any given hypothesis that you want to test, there are several possible methods for performing that test, so you can redo your test many times. This is one type of robustness checking, which includes many other things as well (e.g. running your test over subsamples or resamples of the data, checking for sensitivity to outliers, etc). Good statisticians generally like to do lots and lots of robustness checking.

Re: Psychology Journal Bans Significance Testing

#25
post #20

I spent about a hour explaining p-values to a fellow graduate-level researcher a few weeks ago. I pretty much just kept rephrasing the definition in slightly different ways until the person finally got it. In undergrad, hypothesis testing was more or less taught as "do this inscrutable calculation and if the result is 0.05 or less, you win". The point is, in my experience, a lot of people really don't get p-values, e…

Let me replay it to you and see if I understand it, because I don't know if I do: You assume a null hypothesis that (usually) represents the status quo of no influence between the theory and the data. You then collect data. The p value then describes the probability of that data aligning with / being as a result of the null hypothesis. In other words, a p Is that correct? I may have minced terms there because my stat…

No there's a 5% chance the null hypothesis is correct. the probability your theory is correct is unknowable.

Re: Psychology Journal Bans Significance Testing

#26

>The type of analysis being banned is often called a frequentist analysis I find that there is a trend of associating "bad statistics" with "Frequentists Statistics" which isn't really fair. If you found a statistician trained only in Frequentist methods and asked their opinion on experiment design in psychological research they would likely be just as appalled as any Bayesian. I'm a big fan of Bayesian methods, but…

I think what this journal doing is probably a good thing, but only as the lesser of several evils. The truth is something more like... it is easier for soft sciences to abuse frequentist statistics than bayesian. Both have merit, it cannot be argued, but it is simply easier to produce meaningless conclusions with frequentist statistics done wrong.

This situation is so bad that it merits banning frequentist for this journal and I think that's reasonable. This doesn't mean that every journal in every field should, but perhaps it will be a useful temporary measure to improve quality.

Re: Psychology Journal Bans Significance Testing

#27

There was a question yesterday about Evidence Based Medicine vs Science Based Medicine. The SBM criticism of EBM is the over-reliance on Randomized Controlled Trials that meet p=0.05, without looking at the prior probability that a treatment would help. For example, EMB would say that if you have an RCT that shows that a lucky rabbit's foot works, then you have reasonable evidence to put that into practice. The issue…

> The SBM criticism of EBM is the over-reliance on Randomized Controlled Trials that meet p=0.05, without looking at the prior probability that a treatment would help. The strength of significance testing is that it purposely doesn't try to tell you how likely something is to be true, only how likely the data you got was the result of chance assuming the treatment is no better than placebo. You're still taking the pr…

You're absolutely correct that Bayesian approaches are not magical and do not suddenly supply you with vastly more information than frequentist approaches (particularly when you have a really poor prior, in which case the Bayesian approach will be similarly poor). Bayesian statistics is certainly very popular right now, but it should not be looked at as some sort of panacea for all statistical problems.

However, I would say that Bayesian approaches do have a big advantage in terms of helping with the interpretation problems that plague frequentist significance testing. Namely, as the OP article points out, Bayesian approaches reformulate the testing question in a way that is more intuitive, i.e. "what is the probability of the hypothesis given both the prior probability and the new data?". So yes, Bayesian methods surely do not fix everything, but since interpretation of statistics is such a major concern, they can be quite beneficial.

Re: Psychology Journal Bans Significance Testing

#28
post #20

I spent about a hour explaining p-values to a fellow graduate-level researcher a few weeks ago. I pretty much just kept rephrasing the definition in slightly different ways until the person finally got it. In undergrad, hypothesis testing was more or less taught as "do this inscrutable calculation and if the result is 0.05 or less, you win". The point is, in my experience, a lot of people really don't get p-values, e…

Let me replay it to you and see if I understand it, because I don't know if I do: You assume a null hypothesis that (usually) represents the status quo of no influence between the theory and the data. You then collect data. The p value then describes the probability of that data aligning with / being as a result of the null hypothesis. In other words, a p Is that correct? I may have minced terms there because my stat…

Not quite. The 5% represents the chance that if the null hypothesis is true, you would draw data at least as extreme as the data you just saw in a repeated experiment.

Computing the probability that the data came from the theory stated in the null hypothesis would require a (Baysian) prior.

Also, Tloewald's reply is completely and inexorably wrong. Tloewald seems to want a Bayesian answer, which frequentist statistics can't give you.

Re: Psychology Journal Bans Significance Testing

#29

I spent about a hour explaining p-values to a fellow graduate-level researcher a few weeks ago. I pretty much just kept rephrasing the definition in slightly different ways until the person finally got it. In undergrad, hypothesis testing was more or less taught as "do this inscrutable calculation and if the result is 0.05 or less, you win". The point is, in my experience, a lot of people really don't get p-values, e…

I agree. I'm in the pool of people who don't really understand p-values even after taking 2 statistics courses. The main issue I've noticed is that compared to many other calculations, you just have no idea if your result is correct. You can make wrong assumptions, wrong calculation, apply wrong methods, but in the end you get a number... and that's it. It may be a wrong number and you'll never know. This is complete…

The main issue I've noticed is that compared to many other calculations, you just have no idea if your result is correct. You can make wrong assumptions, wrong calculation, apply wrong methods, but in the end you get a number... and that's it.

The same is true in Bayesian statistics, and even simple formal reasoning with no statistics in sight. If you make wrong assumptions, you'll get the wrong result.

The only thing you can expect statistics to do is help you change your opinion about the relative merits of opposing theories. If both your opposing theories are wrong, you will still be equally wrong.

The true flaw with frequentist statistics is that it goes out of it's way to hide this fact from you. In contrast, Bayesian stats forces you to explicitly choose a prior, enumerate your assumptions, and accept that your conclusion is based on these things.

Re: Psychology Journal Bans Significance Testing

#30
post #20

Earlier quoted context omitted.

Let me replay it to you and see if I understand it, because I don't know if I do: You assume a null hypothesis that (usually) represents the status quo of no influence between the theory and the data. You then collect data. The p value then describes the probability of that data aligning with / being as a result of the null hypothesis. In other words, a p Is that correct? I may have minced terms there because my stat…

Not quite. The 5% represents the chance that if the null hypothesis is true , you would draw data at least as extreme as the data you just saw in a repeated experiment. Computing the probability that the data came from the theory stated in the null hypothesis would require a (Baysian) prior. Also, Tloewald's reply is completely and inexorably wrong. Tloewald seems to want a Bayesian answer, which frequentist statisti…

Thanks, that small distinction does make sense to me. I'm surprised I had it as close as I did.
Post reply on HN