Live data from Hacker News

Psychology Journal Bans Significance Testing

sciencebasedmedicine.org

11–20 of 88 posts

Re: Psychology Journal Bans Significance Testing

#11

There was a question yesterday about Evidence Based Medicine vs Science Based Medicine. The SBM criticism of EBM is the over-reliance on Randomized Controlled Trials that meet p=0.05, without looking at the prior probability that a treatment would help. For example, EMB would say that if you have an RCT that shows that a lucky rabbit's foot works, then you have reasonable evidence to put that into practice. The issue…

> The SBM criticism of EBM is the over-reliance on Randomized Controlled Trials that meet p=0.05, without looking at the prior probability that a treatment would help.

The strength of significance testing is that it purposely doesn't try to tell you how likely something is to be true, only how likely the data you got was the result of chance assuming the treatment is no better than placebo. You're still taking the prior probability into account when trying to figure out the truth, you're just not putting a number on it.

My concern with bayesian approaches is that, like with frequentist approaches, the truth is still fundamentally unknowable, only now you're encouraged to put a number on that and pretend that it's science. While bayesian approaches totally make sense in trying to determine a patient's likelihood of having some disease when there is already data available for the prevalence in a population and the sensitivity and specificity of the tests, using bayesian logic to weight clinical trials strikes me as being highly dubious.

It would be one thing if SBM actually developed a framework to give a weight to each methodological feature of a trial, but so far I haven't seen much work to build a functioning system. Though if you're really honest about all the ways that you can have positive results without something actually being true, it seems like almost no amount of research will ever have a significant effect on the prior.

Re: Psychology Journal Bans Significance Testing

#12
post #8

The problem is this: You need a solid background in statistics and some mathematics to perform these tests properly. Most apparently, people from that field lack these skills. If someone can't hit a nail on the head with a hammer, blaming and banning the use of hammers won't solve a thing.

Taking hammers away will solve the problem of holes in things which aren't supposed to have them, including thumbs, walls, people's heads, and candy jars.

there are just other tools that will get abused. bayes factors for example

Re: Psychology Journal Bans Significance Testing

#13
>The type of analysis being banned is often called a frequentist analysis

I find that there is a trend of associating "bad statistics" with "Frequentists Statistics" which isn't really fair. If you found a statistician trained only in Frequentist methods and asked their opinion on experiment design in psychological research they would likely be just as appalled as any Bayesian.

I'm a big fan of Bayesian methods, but the solutions of "we'll solve the problem of misunderstanding p-values by removing them!" is still a problem of misunderstanding p-values! The misunderstanding is the issue, not the p-value.

Re: Psychology Journal Bans Significance Testing

#14
post #4

I wonder if it just boils down to this: Exclusive reliance on any single tool by an entire field for a long enough time period will eventually lead to a proliferation of bad results.

More like exclusive reliance on a single tool will lead to a requirement for everybody to use it, even if they don't understand how to use it properly, which leads to both unintentional misuse (by experimenters who don't understand it) and the inability to catch intentional misuse (because adjudicators don't understand it either).

Re: Psychology Journal Bans Significance Testing

#15
I spent about a hour explaining p-values to a fellow graduate-level researcher a few weeks ago. I pretty much just kept rephrasing the definition in slightly different ways until the person finally got it. In undergrad, hypothesis testing was more or less taught as "do this inscrutable calculation and if the result is 0.05 or less, you win". The point is, in my experience, a lot of people really don't get p-values, even graduate-level and post-graduate scientists.

Re: Psychology Journal Bans Significance Testing

#16
Here's the key passage:

BASP will require strong descriptive statistics, including effect sizes. We also encourage the presentation of frequency or distributional data when this is feasible. Finally, we encourage the use of larger sample sizes than is typical in much psychology research, because as the sample size increases, descriptive statistics become increasingly stable and sampling error is less of a problem.

In other words: report the effect size, plot the data, and increase the N.

Re: Psychology Journal Bans Significance Testing

#17
post #8

The problem is this: You need a solid background in statistics and some mathematics to perform these tests properly. Most apparently, people from that field lack these skills. If someone can't hit a nail on the head with a hammer, blaming and banning the use of hammers won't solve a thing.

Depends what are you trying to solve - people's understanding of p-values, or many false results? If you care about getting true results and there are easier ways to get there than p-values, why would it be a bad solution?

Similar thing to what happens in software security. Sure, you could tech people how to use memory management properly - and then watch them fail again and again. Or you could just provide an alternative solution like rust which changes the issue completely.

Re: Psychology Journal Bans Significance Testing

#18
post #7

The article is certainly correct that p-values and confidence intervals (or confidence sets, in multi-dimensional contexts) are widely misunderstood, not just in psychology or other social sciences, but in the hard sciences as well. The problem is even worse when you look outside of academia at common practices in more applied settings. As suggested, a good approach is to take p-values not as conclusive or decisive,…

[deleted]

Re: Psychology Journal Bans Significance Testing

#19

I spent about a hour explaining p-values to a fellow graduate-level researcher a few weeks ago. I pretty much just kept rephrasing the definition in slightly different ways until the person finally got it. In undergrad, hypothesis testing was more or less taught as "do this inscrutable calculation and if the result is 0.05 or less, you win". The point is, in my experience, a lot of people really don't get p-values, e…

I agree. I'm in the pool of people who don't really understand p-values even after taking 2 statistics courses. The main issue I've noticed is that compared to many other calculations, you just have no idea if your result is correct. You can make wrong assumptions, wrong calculation, apply wrong methods, but in the end you get a number... and that's it. It may be a wrong number and you'll never know. This is completely different to many other practical applications of math where you can verify your result in various ways, or validate your answer against the initial assumptions, or test your program against lots of inputs.

Re: Psychology Journal Bans Significance Testing

#20

I spent about a hour explaining p-values to a fellow graduate-level researcher a few weeks ago. I pretty much just kept rephrasing the definition in slightly different ways until the person finally got it. In undergrad, hypothesis testing was more or less taught as "do this inscrutable calculation and if the result is 0.05 or less, you win". The point is, in my experience, a lot of people really don't get p-values, e…

Let me replay it to you and see if I understand it, because I don't know if I do:

You assume a null hypothesis that (usually) represents the status quo of no influence between the theory and the data. You then collect data. The p value then describes the probability of that data aligning with / being as a result of the null hypothesis. In other words, a p Is that correct? I may have minced terms there because my stats training is woefully inadequate, but I think I adequately conveyed the concept?

Post reply on HN