Live data from Hacker News

Is it time to up the statistical standard for scientific results?

arstechnica.com

11–20 of 44 posts

Re: Is it time to up the statistical standard for scientific results?

#11
The problem with upping the standard is, as the PNAS article acknowledges, that you need a larger sample size to produce any given result.

Unfortunately, many studies -- particularly those in medicine -- are already conducted with samples that are too small to detect any effect you'd reasonably expect to see, because many researchers do not calculate in advance what sample size would be required. This has interesting paradoxical effects: the only published studies are those that overestimate the size of the true effect.

http://www.refsmmat.com/statistics/power.html http://www.refsmmat.com/statistics/regression.html#truth-inf...

So there's a tradeoff. Do you want to eliminate false positives at the cost of more false negatives? It's a difficult balance. I suspect there are many areas where poor statistical practice can be remedied to produce better results without greater expense.

Re: Is it time to up the statistical standard for scientific results?

#12
post #7
post #2

> These are the sorts of nuts-and-bolts reproducibility issues that drive researchers crazy, because they can be affected by things like the specific strain of mice you use, where you buy your chemicals, and even the pH of your lab's water supply. No amount of statistical thinking is going to change any of that. This screams systematic error and error propagation to me. It's possible that we don't need to up the p-va…

exactly wrong, it's the multidisciplinary scientists who don't get enough depth or rigor to come close to understanding sources of error and error propagation. The thing I fear is underexperienced multidisciplinary researchers coming up with convoluted procedures that put their hands and eyes on their work in so many places that small sources of error have compounded themselves and entrenched themselves into nooks an…

You're right! It's those other jerks getting all their high school statistics wrong. No true scientist (Scotsman) could ever make THAT mistake.

Re: Is it time to up the statistical standard for scientific results?

#13
The real problem is that effect size is rarely discussed. P-values only relate to variability, sometimes high variability is acceptable, sometimes it's not. Taken by itself, a p-value is worthless, no matter how small it is. For example, suppose you develop a fertilizer that you're 99.99999% certain will produce one additional ear of corn in 10,000 bushels. Who cares!? You see a lot of that in published papers. A lot of researchers stop once they get a p-value below 0.05 and neglecting effect size is the norm.

Re: Is it time to up the statistical standard for scientific results?

#14

The problem with upping the standard is, as the PNAS article acknowledges, that you need a larger sample size to produce any given result. Unfortunately, many studies -- particularly those in medicine -- are already conducted with samples that are too small to detect any effect you'd reasonably expect to see, because many researchers do not calculate in advance what sample size would be required. This has interesting…

I still think the answer is not in any particular standard in any particular field, rather it's outputting the final, worked datasets. Making the numbers themselves available and usable to the 'reader' or scientist following up will immediately clear up issues of reproducibility or not meeting statistical thresholds. If I see the same data, run it through my own processes and it looks like noise I'm likely to discount the results moving forward. There's nothing inherently wrong with saying 'look, I did this once, and this is what I saw'. And many times it's useful. It's less useful, but still reasonable in certain circumstances to say, 'look, I did this a hundred times and I saw this once'. There are occasions when the experiment cannot reasonably have more than a statistically insignificant 'n' - but it still might be useful to see what happens. What is not reasonable is to say, 'this is what I (sometimes) see always (trust me)' and hide your data behind a 'representative' jpg and a 'statistically significant' p-value.

As I scientist who has worked with both rich, and poor datasets from standard and invented datatypes, I'd just say, "let me see the data".

Re: Is it time to up the statistical standard for scientific results?

#15
post #7
post #2

> These are the sorts of nuts-and-bolts reproducibility issues that drive researchers crazy, because they can be affected by things like the specific strain of mice you use, where you buy your chemicals, and even the pH of your lab's water supply. No amount of statistical thinking is going to change any of that. This screams systematic error and error propagation to me. It's possible that we don't need to up the p-va…

exactly wrong, it's the multidisciplinary scientists who don't get enough depth or rigor to come close to understanding sources of error and error propagation. The thing I fear is underexperienced multidisciplinary researchers coming up with convoluted procedures that put their hands and eyes on their work in so many places that small sources of error have compounded themselves and entrenched themselves into nooks an…

I haven't seen any reason to believe that being multidisciplinary makes this more likely. The whole problem is caused by the fact that most of the single-disciplinary researchers are using bad statistics.

Very few degree programs give scientists enough background in statistics and experimental rigor to prevent these problems.

Re: Is it time to up the statistical standard for scientific results?

#16

>In most fields, if there's less than a five percent chance that you'd get the two numbers by random chance, then you can reject chance—the results are considered significant. In statistical terms, this is called having a p value of less than 0.05. No it isn't [1]. [1]: https://en.wikipedia.org/wiki/P-value

The Ars definition of a p value is not precisely worded but, I think, reasonably accurate. It doesn't include the bit about also including the probability of obtaining results more extreme than what you obtained. The best definition I know is > The P value is defined as the probability, under the assumption of no effect or no difference (the null hypothesis), of obtaining a result equal to or more extreme than what w…

Yeah that definition is good.

For the purposes of understanding the definition, another way of looking at it is basically a 'statistical proof by contradiction':

1. Assume null hypothesis is true

2. Compute test statistic

3. Ask the question, "What is the probability of obtaining that test statistic or one more extreme?" (this probability is the p-value)

4. Pick a threshold (usually 0.05 but this is totally arbitrary)

5. If p Reductio ad statistico absurdum.

Re: Is it time to up the statistical standard for scientific results?

#17
In most areas (social sciences, medicine) there is no "standard" for statistical significance. In economics, people try to deal with the issue by looking at the robustness of a result. If a result remains when you make various changes to the specification of your model, then it is less likely to be a statistical artifact.

Another step in the right direction is greater reproducibility, so that at least people can play with your analysis and see if what you did was the most direct and natural analysis, or if there are clear signs of playing with the parameters until you get the result you want.

Re: Is it time to up the statistical standard for scientific results?

#19
post #7

Earlier quoted context omitted.

exactly wrong, it's the multidisciplinary scientists who don't get enough depth or rigor to come close to understanding sources of error and error propagation. The thing I fear is underexperienced multidisciplinary researchers coming up with convoluted procedures that put their hands and eyes on their work in so many places that small sources of error have compounded themselves and entrenched themselves into nooks an…

I haven't seen any reason to believe that being multidisciplinary makes this more likely. The whole problem is caused by the fact that most of the single-disciplinary researchers are using bad statistics. Very few degree programs give scientists enough background in statistics and experimental rigor to prevent these problems.

no doubt, the single-disciplinary researchers use bad statistics, too. My broader point is that multidisciplinary researchers are less likely to use statistics correctly, if for no other reason than that the fact that they are multidisciplinary means that they are less likely to be able to handle depth from a personality/self-habits point of view.

Perhaps that's a little bit of projection, but I don't use any but the most rudimentary statistics in my publications and I don't make claims about significance.

Re: Is it time to up the statistical standard for scientific results?

#20

>In most fields, if there's less than a five percent chance that you'd get the two numbers by random chance, then you can reject chance—the results are considered significant. In statistical terms, this is called having a p value of less than 0.05. No it isn't [1]. [1]: https://en.wikipedia.org/wiki/P-value

Indeed. And any statistician worth their salt doesn't reject chance, they fail to reject the null hypothesis.
Post reply on HN