Live data from Hacker News

Scientists Perturbed by Loss of Stat Tool to Sift Research Fudge from Fact

scientificamerican.com

21–30 of 40 posts

Re: Scientists Perturbed by Loss of Stat Tool to Sift Research Fudge from Fact

#21
post #4

Here's the editorial by David Trafimow and Michael Marks explaining the new policy for their journal "Basic and Applied Social Psychology": http://www.tandfonline.com/doi/pdf/10.1080/01973533.2015.101... And here's their concluding paragraph explaining their rationale and objective: We conclude with one last thought. Some might view the NHSTP[1] ban as indicating that it will be easier to publish in BASP, or that les…

And what should be used as an alternative? There is no reason to believe banning p value would result in improved research quality. Desperate researchers will just find another technique to game.

Instead of blindly applying a tool you don't understand, you'll need to build a statistical model, explain why it's valid, and then construct a meaningful measurement based on it.

Then the referee and reader will be required to understand it and will have the ability to critique it.

Re: Scientists Perturbed by Loss of Stat Tool to Sift Research Fudge from Fact

#22
I find this whole backlash against p-values pretty confusing. That is probably because I come from particle physics, where we also use a lot of statistics, but in subtly different ways.

Hypothesis testing is not too hard [1]. You pick a cutoff, say p Since we are a cautious bunch, we actually put the threshold for discovery at 5 sigma - pOne other thing that we have to take into account - and many people forget this - is the look-elsewhere effect. If you perform one search, looking for e.g. a Higgs Boson with a mass of 126 GeV, you expect N events in your experiment if it is not there, and N+X if it is there. You know how N is distributed, and the interpretation is straightforward. However if you perform a scan, looking at 120, 121, 122, 123... GeV, then you have to adjust your p-value, since you are basically performing a bunch of different experiments, and by chance alone some of them are bound to turn up "significant".

The same thing applies when hundreds or thousands of Master and PhD students and postdocs do their analyses - even if no one makes a mistake, some of them will "find" a 3 sigma or larger effect that isn't there, just due to the sheer number of independent statistical tests performed. I've "found" new particles myself this way, but when you keep calm, put it into context by looking at other analyses, and try to add more data, you'll often find that your result melts away.

------------

[1] explaining it is hard, and I will undoubtedly have messed up, especially since I'm tired.

Re: Scientists Perturbed by Loss of Stat Tool to Sift Research Fudge from Fact

#23
>>Several journals are trying a new approach...in which researchers publicly “preregister” all their study analysis plans in advance. This gives them less wiggle room to engage in the sort of unconscious—or even deliberate—p-hacking that happens when researchers change their analyses in midstream to yield results that are more statistically significant than they would be otherwise. In exchange, researchers get priority for publishing the results of these preregistered studies—even if they end up with a p-value that falls short of the normal publishable standard.

It's not exactly the same issue as the one addressed by banning p-values, but this would help a lot.

Re: Scientists Perturbed by Loss of Stat Tool to Sift Research Fudge from Fact

#24

I find this whole backlash against p-values pretty confusing. That is probably because I come from particle physics, where we also use a lot of statistics, but in subtly different ways. Hypothesis testing is not too hard [1]. You pick a cutoff, say p Since we are a cautious bunch, we actually put the threshold for discovery at 5 sigma - p One other thing that we have to take into account - and many people forget this…

> we actually put the threshold for discovery at 5 sigma - pThis says more about the sillyness of statisticians than that of particle physicists.

Re: Scientists Perturbed by Loss of Stat Tool to Sift Research Fudge from Fact

#25

I find this whole backlash against p-values pretty confusing. That is probably because I come from particle physics, where we also use a lot of statistics, but in subtly different ways. Hypothesis testing is not too hard [1]. You pick a cutoff, say p Since we are a cautious bunch, we actually put the threshold for discovery at 5 sigma - p One other thing that we have to take into account - and many people forget this…

> I find this whole backlash against p-values pretty confusing.

Plenty of research (not in particle physics) uses the cut-off of 0.05 instead, misunderstands the p-value as "probability that the result is not real", and also ignores the fact that a p-value of 0.05 can too often be reached by making a variety of choices in experimental setup and data analysis. When the headline calls it a "Tool to Sift Research Fudge from Fact", NHST wasn't doing this job well at all, and wasn't really designed for this job in the first place. There are far, far too many people who seem to think that p<0.05 means something is a "research fact", which is a failure of statistics education as well.

Re: Scientists Perturbed by Loss of Stat Tool to Sift Research Fudge from Fact

#26

I find this whole backlash against p-values pretty confusing. That is probably because I come from particle physics, where we also use a lot of statistics, but in subtly different ways. Hypothesis testing is not too hard [1]. You pick a cutoff, say p Since we are a cautious bunch, we actually put the threshold for discovery at 5 sigma - p One other thing that we have to take into account - and many people forget this…

> which sometimes gets us ridiculed by statisticians

You get laughed at because you have this huge luxury at dissecting huge amounts of data, running the experiment billions of times.

Do you understand that in most disciplines, you don't have that luxury?

Re: Scientists Perturbed by Loss of Stat Tool to Sift Research Fudge from Fact

#27
post #15
post #12

Earlier quoted context omitted.

At one top British university, the first year Physics practicals are universally reviled, everyone gets within a few marks of each other, the experiments are trivial (roll a ball down a slope!), and at every meeting the academics all agree that they'd prefer that they were dropped. Unfortunately there's some kind of government requirement for practical work, so they stay.

>the experiments are trivial (roll a ball down a slope!) they should replace it with trivial experiments of the 21st century - photon counting for Bell inequalities testing. Wrt. the original article - good riddance, one less orthodoxy in science. Because it isn't about a tool - p-values in this case - itself, it is about orthodoxy which is the main enemy of science.

Oh exactly, sounds like there is ways of getting around the government mandate.

Re: Scientists Perturbed by Loss of Stat Tool to Sift Research Fudge from Fact

#28

I find this whole backlash against p-values pretty confusing. That is probably because I come from particle physics, where we also use a lot of statistics, but in subtly different ways. Hypothesis testing is not too hard [1]. You pick a cutoff, say p Since we are a cautious bunch, we actually put the threshold for discovery at 5 sigma - p One other thing that we have to take into account - and many people forget this…

The interesting problem is that if you set a strict significance threshold but have a sample size that is too small, you will still sometimes correctly get significant results, but the effect sizes will all be exaggerated.

If the sample size is too low, the only effect size big enough to be significant may be one that is much larger than the truth. So you only claim significance when you get lucky and exaggerate.

This is actually an enormous problem, and it probably affects physics too. Many medical and biological papers tout huge effects which turn out to be completely wrong -- not because their result was a false positive but because it's an exaggerated true positive. Sample sizes tend to be hilariously inadequate in soft sciences where each data point costs thousands of dollars.

I call this "truth inflation," although I don't know if it's been discussed enough to have a common name. It's heavily discussed in my book, if you're interested in seeing why the backlash against p values is so widespread: http://www.statisticsdonewrong.com/regression.html#truth-inf...

Re: Scientists Perturbed by Loss of Stat Tool to Sift Research Fudge from Fact

#29
I think it's easier to explain this in terms of likelihood theory. Likelihood is the probability of observed data GIVEN A SPECIFIC MODEL. This is NOT to be confused with the probability of a specific model being correct GIVEN THE OBSERVED DATA. It is the later probability that people want to really know, but confusing it with the former probability can have catastrophic consequences in fields like medicine, engineering, jurisprudence, finance and insurance.

The problem is related to the Prosecutor's Fallacy (https://en.wikipedia.org/wiki/Prosecutor%27s_fallacy): "Consider this case: a lottery winner is accused of cheating, based on the improbability of winning. At the trial, the prosecutor calculates the (very small) probability of winning the lottery without cheating and argues that this is the chance of innocence. The logical flaw is that the prosecutor has failed to account for the large number of people who play the lottery."

There is a mathematical statistics professor at the University of Toronto named D.A.S. Fraser (http://www.utstat.utoronto.ca/dfraser/) who is an expert in likelihood theory and has commented on this issue: "...statistics does have the answer! The answer is contained in the p-value function from likelihood theory: p(delta) Here delta is the relevant parameter with delta_0 as the null value and delta_1 as the alternative needing detection. Then p(delta_0) is the observed p-value, p(delta_1) is the detection probability, and the rest is judgement: the route to the Higgs boson."

Re: Scientists Perturbed by Loss of Stat Tool to Sift Research Fudge from Fact

#30
Their issue is basically with the lack of negative result reporting that makes p-values useless; it seems very odd to vilify a valuable tool when it's never going to solve the real problem that people generally re-run experiments until they stumble upon a publishable metic.
Post reply on HN