Live data from Hacker News

Psychology Journal Bans Significance Testing

sciencebasedmedicine.org

61–70 of 88 posts

Re: Psychology Journal Bans Significance Testing

#61

Earlier quoted context omitted.

You're still taking the prior probability into account when trying to figure out the truth, you're just not putting a number on it. This is a danger sign - you are doing the same things the Bayesians do, just informally, less explicitly, and probably incorrectly. The fact is that to make a good decision, eventually you need to compute a single number. This is an elementary fact of topology: https://www.chrisstucchio.…

> That number will be based on some unproveable assumptions. Given that, would you support using a random number generator as part of the drug approval process to remind people of the importance of the unknown and unknowable?

[edit: I previously said I didn't understand Alex's point.] I now understand the point you were trying to make. A better way to put it - if a random number generator were used in a decision process, I'd favor making the algorithm and random seed explicit.

Any procedure you use will have assumptions. You can't escape this. The only question is whether we show or hide them. Can you give an argument in favor of hidden assumptions and non-explicit procedures?

Re: Psychology Journal Bans Significance Testing

#62
I might be missing the forest from the tree's.

The paper talks about how it seems researches are "hacking" (their word) p-values.If the researchers lack the ethics using one form of statistics, what is really stopping them from misusing Bayesian Analysis?

Sidenote: While we are talking about bayesian stuff. I recently ran into the sleeping beauty problem ( http://en.wikipedia.org/wiki/Sleeping_Beauty_problem ) and it showed how different interpretations of an experiment can lead to people believing in two very different answers in an almost religious way. Some people argued this thought experiment shows a clear flaw in bayesian thinking. I'm not sure either way, I'm still thinking about the puzzle.

Re: Psychology Journal Bans Significance Testing

#63

>The type of analysis being banned is often called a frequentist analysis I find that there is a trend of associating "bad statistics" with "Frequentists Statistics" which isn't really fair. If you found a statistician trained only in Frequentist methods and asked their opinion on experiment design in psychological research they would likely be just as appalled as any Bayesian. I'm a big fan of Bayesian methods, but…

I think the problem is not "if you find a frequentist (as opposed to bayesian) statistician", but "if you find a frequentist (as opposed to bayesian) e.g. biologist".

Non-statisticians have been trained using bad, frequentist methods, and one way of forcing them to retrain is by forcing them to learn new statistical tools to get published.

Re: Psychology Journal Bans Significance Testing

#64
post #55

Earlier quoted context omitted.

You have heard "19 times out of 20" described in the news? That is the 0.05 restated for laypeople. 1 time out of 20 you will get a false positive, in this case that the rabbit's foot worked.

Sorry maybe I'm being dense, but who would take 1 out of 20 success to mean they should start buying rabbit feet?

Nobody. The problem is that if you conduct 20 trials of the efficacy of rabbit feet, you'd expect 1 trial to show a significant effect. (if you're interpreting P-values the way many people do, incorrectly.)

Re: Psychology Journal Bans Significance Testing

#65
post #55

Earlier quoted context omitted.

You have heard "19 times out of 20" described in the news? That is the 0.05 restated for laypeople. 1 time out of 20 you will get a false positive, in this case that the rabbit's foot worked.

Sorry maybe I'm being dense, but who would take 1 out of 20 success to mean they should start buying rabbit feet?

When you do a single trial you might show that rabbits feet are effective.

You don't yet know if this is a 1 in 20 result or a 19 in 20 result.

That's why you replicate.

Re: Psychology Journal Bans Significance Testing

#66

If the stats aren't based on samples of real-world data, rather than very rigorously designed experimental comparisons, I think there's another reason to be wary of p-values. [Note, in what follows I probably use some terminology wrong, because I'm not a statistician, but I do think the point is important and I don't see much written about it.] In the real world data is not a bunch of independent events, but events (…

Could you expand on what you mean by "interconnectedness within website traffic"? If one person visiting the site has no influence on other people visiting the site, then measurements of their behavior will be independent. If Facebook tests a different interface on half of their users and the changed behavior of those users indirectly has some impact on the behavior of the control group, then your measurements would…

The presenter I mentioned did not go into details about what interconnectedness he found, but I think it's quite obvious that people visiting a site do have influence on other visitors, which is at least a part of the underlying issue. On the simplest level, most web sites have share buttons to make it as easy as possible for visitors to influence other traffic. Or, other examples, a trending tweet can massively influence patterns of usage of a website or Facebook page (or many tweets with small reach can cause many small influences); or an RSS feed might influence patterns of tweeting or posting elsewhere. I think there are myriad other interconnections within web traffic. We do a great deal of work to drive traffic that is based on the premise that different visitors to websites are mostly interconnected. It's these factors that give me pause when I think about statistical measures that are premised on an assumption that we're measuring independent events.

Re: Psychology Journal Bans Significance Testing

#67

Earlier quoted context omitted.

> That number will be based on some unproveable assumptions. Given that, would you support using a random number generator as part of the drug approval process to remind people of the importance of the unknown and unknowable?

[edit: I previously said I didn't understand Alex's point.] I now understand the point you were trying to make. A better way to put it - if a random number generator were used in a decision process, I'd favor making the algorithm and random seed explicit. Any procedure you use will have assumptions. You can't escape this. The only question is whether we show or hide them. Can you give an argument in favor of hidden a…

> Can you give an argument in favor of hidden assumptions and non-explicit procedures?

So as counterintuitive as it sounds, I think there are actually a couple of good arguments that can be made here:

1) With significance testing, the burden of supplying the assumptions and determining meaning is largely on the reader. With bayesian, it's transferred to the author. While it might make sense to use Bayesian for things like the Cochrane report, it's not obvious to me that each person who designs a research study and collects/analyzes data should also be in the business of trying to say whether some phenomena is real when looking at all other studies.

Essentially each study now becomes a metastudy, with all of the practical and epistemological problems that entails. The fact that it's difficult to figure out what that even means should be a red flag. (And yes, I realize this is the Chewbacca defense.)

2) So TokenAdult actually turned me onto this book Measurement In Psychology, which is all about the epistemological problems with assuming that anything you can assign a number to is a measurement. That is, having the property of being meaningful when interpreted on a ratio scale. The exact argument is kind of esoteric, but the basic takeaway is that it's very easy to trick yourself into thinking that just because you can assign a number to something that it's a measurement, to the point where assigning numbers to things in the first place tends to lead to worse decision making than if you had just used a green/yellow/red system or whatever.

Re: Psychology Journal Bans Significance Testing

#68
the main reason to ban significance testing in all fields is because of logicallee's getcher scientific results agency.

Our prices are:

$10,000 random study, no results guaranteed. Not recommended! Highly likely to be damaging.

$20,000 basic study, p=0.5 - study inconclusive but implies it's at least not "more likely" that the damaging/negative result is correct. (Assuming uniform priors.) No scientific value.

$100,000 weak FUD. Suggests that the damaging/negative results (assuming bayesian reasoning or uniform priors) may be incorrect, at a suggestive p$200,000 basic refutation. Refutes the damaging/negative result at a statistically significant p$500,000 Silver refutation. Highly significant refutation at p$1,000,000 Gold refutation. Highly significant refutation at pACADEMIC BONUS: prices are free for tenured professors, thousands of whom can do all the studies they want and only publish if they see some significant effect.

EDIT:

in other words, http://xkcd.com/882/

Re: Psychology Journal Bans Significance Testing

#69

I spent about a hour explaining p-values to a fellow graduate-level researcher a few weeks ago. I pretty much just kept rephrasing the definition in slightly different ways until the person finally got it. In undergrad, hypothesis testing was more or less taught as "do this inscrutable calculation and if the result is 0.05 or less, you win". The point is, in my experience, a lot of people really don't get p-values, e…

> I pretty much just kept rephrasing the definition in slightly different ways until the person finally got it

well don't keep us hanging... (what was it?)

Re: Psychology Journal Bans Significance Testing

#70

Earlier quoted context omitted.

Could you expand on what you mean by "interconnectedness within website traffic"? If one person visiting the site has no influence on other people visiting the site, then measurements of their behavior will be independent. If Facebook tests a different interface on half of their users and the changed behavior of those users indirectly has some impact on the behavior of the control group, then your measurements would…

The presenter I mentioned did not go into details about what interconnectedness he found, but I think it's quite obvious that people visiting a site do have influence on other visitors, which is at least a part of the underlying issue. On the simplest level, most web sites have share buttons to make it as easy as possible for visitors to influence other traffic. Or, other examples, a trending tweet can massively infl…

Hmmm, again, there are certainly ways that interactions between visitors can cause statistical dependence, but not in the specific case you mention. Let's take an A/B test on a referral funnel. If a user invites all of his friends, and his friends then visit the site, they will be randomized over A and B just like the original user, and so any effect that is not due to changes in the referral experience will simply not matter because it will contribute equally to both groups.

Without better examples it's very hard to judge whether this is a real problem.

Post reply on HN