Live data from Hacker News

Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

fivethirtyeight.com

121–130 of 130 posts

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#121
post #58

Earlier quoted context omitted.

Thanks for the very interesting thoughts. Could you explain a bit how your rule of thumb works and why it's better than p-values? Why is the difference vs. the square root of the max available sample size a meaningful measure?

The idea is that you will decide when you've either expended as much energy as you are willing to, or when you're convinced that you won't make a different decision. There is a simple symmetry argument that shows that the odds of a random walk reaching sqrt(N) in one direction and then getting to the opposite direction by N observations is exactly the same as the odds of a random walk reaching 2 sqrt(N) by the time y…

Did you know that your rule of "do the tests until significance is reached" denies the use of chain probability and most tests based on it?

You can show that in the limit, your rule will always reach any given threshold just by pure luck.

You really need to set the number of trials or samples before performing them.

Here's a more in depth explanation: http://www.evanmiller.org/how-not-to-run-an-ab-test.html

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#122
post #69

Earlier quoted context omitted.

One issue is that if you have a large effect that's consistently and easily reproduced, you don't actually need very accurate measurements or a statistical analysis at all. So any minimum standard would need to take into account. Another issue is that science is expensive and we need to make decisions all the time whether there is any science backing them or not. So what do you do if there's no science that meets the…

> One issue is that if you have a large effect that's consistently and easily reproduced, you don't actually need very accurate measurements or a statistical analysis at all. I mostly agree with this, but you never have that in a difficult decision. > Another issue is that science is expensive and we need to make decisions all the time whether there is any science backing them or not. So what do you do if there's no…

>> One issue is that if you have a large effect that's consistently and easily reproduced, you don't actually need very accurate measurements or a statistical analysis at all.

> I mostly agree with this, but you never have that in a difficult decision.

Consider the problem "should we cure congenital deafness in infants?"

To Deaf Community Leaders, this isn't a difficult decision at all; they are strongly against curing deaf children because it makes their power base smaller.

To deaf parents of deaf children, this is a tricky choice. Curing the child's deafness dramatically improves its prospects in life, but it also inevitably cuts the child off from the parents' community. You have to choose between how good you want your child's future to be, and how close you'd like your relationship with them to be.

To hearing parents of deaf children, this is again a trivial choice; obviously you'd cure the child.

BUT, curing deafness is definitely (1) a large effect that is (2) easily reproduced.

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#124
This is one of the best conversation threads I've ever seen on HN. It's both polite and informative.

I want to toss in my own thoughts here.

Since I've spent the vast majority of my tech career in the Market Research industry (hello, bias!), I'm tempted to say that one of the most frequent intersections between statistical science and business decisions happens in that world.

Product testing, shopper marketing, A/B testing . . . these are pretty common fare these days. But I feel like the MR people are sort of their own worst enemy in many cases.

It's a fairly recent development that MR people are even allowed a seat at the table for major product or business decisions. And when the data nerds show up at the meeting, we have to make human communication decisions that are difficult.

I can't show up at the C-suite and lecture company executives about the finer points of statistical philosophy. When I'm presenting findings to stake-holders, it's my job to abstract the details and present something that makes a coherent case for a decision, based on the data we have available.

It is sinfully attractive to go tell your boss's boss's boss that we have a threshold--a number we can point to. If this number turns out to be smaller than .05, this project is a go.

Three months later, you go back to that boss and tell him the number came back and it was .0499999. The boss says, "Okay, go!" And then you are all, "Wait, wait, wait. Hang on a second. Let's talk about this."

My god, what have I done?

The practical reality of the intersection of statistics and business is a harsh one. We have to do better. In terms of leaky abstractions, the communication of data science to business decision makers is quite possibly the leaky-est of all.

Why is it so leaky? I have two points about this.

1) Statistics is one of the most existentially depressing fields of study. There is no acceptance; there is no love; there is nothing positive about it. Ever.

Statistics is always about rejection and failure. We never accept or affirm a hypothesis. We only ever reject the null hypothesis or we fail to reject it. That's it.

2) In business, we tend to be very very sloppy about formulating our hypotheses. Sometimes we don't even really think about them at all.

Take a common case for market research. New product testing. We do a rep sample with a decent size (say, 1800 potential product buyers) and we randomly show five different products, one of which is the product the person already owns/uses (because that's called control /s). The other 4 products are variations on a theme with different attributes.

What's the null hypothesis here? Does it ever get discussed?

What's the alternative hypothesis?

The implicit and never-talked-about null is that all things being equal, there is no difference between the distribution of purchase likelihood among all products. The alternative is that there is a real difference on a scale of likely to purchase.

The implicit and intuitive assumption is that there is something about that feature set that drives the difference. (I'm looking at you, Max Diff)

But that's not real. It's not a part of the test. The only test you can do in that situation is to check if those aggregate distributions are different from each other. The real null is that they are the same, and the alternative is that they are different.

All you can do with statistics is tell if two distributions are isomorphic.

Now, who wants to try to explain any of that to your CEO? No one does. Your CEO doesn't want it, you don't want it, your girlfriend doesn't want it. No one wants it.

So we try to abstract, and I feel like we mostly fail at doing a good job of that.

This is getting really long, and I don't want to rant. So to finish up, an idea for more effective uses of data science as it interacts with the business world:

I agree, let's stop talking about p values. Let's work harder and funnel the results of those MR studies into practical models of the business' future. Let's take the research and pipe it into Bayesian expected value models.

Let's stop showing stacked bar charts to execs and expecting them to make good decisions based on weak evidence we got from hypotheses we didn't really think about in the first place.

Some of this might come across as a rant. I hope it is not taken that way. This is a real problem that I've been thinking about for a long time. And I don't mean to step on anyone's toes. I have certainly committed many of the data sins that I'm deriding above.

Edited to add:

The real workings of statistics are unintuitive. I'm not saying that they are wrong. But in working with people for years now, I understand the confusion. It's a psychological problem. Hypotheses are either not really well though out or not considered in an organized way, in my experience.

A hypothesis is not concrete in many practical cases. It's a thought. An idea, perhaps. It's often a thing that floats around in your mind, or maybe you paid some lip service and tossed it into your note-taking app.

Data seem much more real. You download a few gigabytes of data and start working on it. It's quite easy to get confused.

I have real data! This is tangible stuff. Thinking of things properly and evaluating the probability of your data given the hypothesis is hard. Your data seems much more concrete. These are real people answering real questions about X.

Even for people who are really hell-bent on statistical rigor, this is a challenge.

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#125
post #69

Earlier quoted context omitted.

> One issue is that if you have a large effect that's consistently and easily reproduced, you don't actually need very accurate measurements or a statistical analysis at all. I mostly agree with this, but you never have that in a difficult decision. > Another issue is that science is expensive and we need to make decisions all the time whether there is any science backing them or not. So what do you do if there's no…

>> One issue is that if you have a large effect that's consistently and easily reproduced, you don't actually need very accurate measurements or a statistical analysis at all. > I mostly agree with this, but you never have that in a difficult decision. Consider the problem "should we cure congenital deafness in infants?" To Deaf Community Leaders, this isn't a difficult decision at all; they are strongly against curi…

Having thought a bit more about this, I want to disagree more strongly with the sentiment "you never have [large, reproducible effects] in a difficult decision". I think difficult decisions necessarily involve that kind of effect.

A couple of things might make a decision difficult:

- Some course of action will produce a big effect. Would that be good or bad?

-- Ok, assume there's a good effect out there to be achieved. Is it worth the cost of obtaining it?

Those questions, variously applied, occupy a lot of people's time and brainpower. In the curing-a-deaf-child example, the parents are making a tradeoff between quality of their child's life (which is good), and closeness to their child (also good). Paying one for the other means giving up something good, which makes the decision nontrivial. But this is a difficult decision because the effects are large, not because they're small.

In contrast, if you're dealing with an effect of very small size, or one that can't be reproduced ("lose weight by following our new diet!")... this might feel like a difficult decision, but it shouldn't. It does not matter what you decide, because (by hypothesis!) your decision will have no effect on the outcome! (Or, for the "very small effect size" case, at most a very small effect.)

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#126
The #1 problem is with p-values is the word "significant". We should use "detectable" instead. Significant implies meaningful to most people, but not in a statistical context. This is quite confusing. Detectable is better because the mainstream meaning aligns with the jargon.

So:

> "Discovering statistically significant biclusters in gene expression data"

becomes:

> "Discovering statistically detectable biclusters in gene expression data"

This rephrasing makes it evident that "statistically detectable" adds little to the title. So the title becomes

> "Discovering biclusters in gene expression data"

A better title.

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#127
This excellent and amusing article by Gerd Gigerenzer discusses the history of p-values and their (mis)use:

"Mindless Statistics"

http://library.mpib-berlin.mpg.de/ft/gg/GG_Mindless_2004.pdf

or

http://www.unh.edu/halelab/BIOL933/papers/2004_Gigerenzer_JS...

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#128
post #58

Earlier quoted context omitted.

The idea is that you will decide when you've either expended as much energy as you are willing to, or when you're convinced that you won't make a different decision. There is a simple symmetry argument that shows that the odds of a random walk reaching sqrt(N) in one direction and then getting to the opposite direction by N observations is exactly the same as the odds of a random walk reaching 2 sqrt(N) by the time y…

Did you know that your rule of "do the tests until significance is reached" denies the use of chain probability and most tests based on it? You can show that in the limit, your rule will always reach any given threshold just by pure luck. You really need to set the number of trials or samples before performing them. Here's a more in depth explanation: http://www.evanmiller.org/how-not-to-run-an-ab-test.html

The procedure I suggested starts with setting the maximum number of conversions before beginning the test. It is emphatically not the procedure that Evan Miller criticizes.

In fact see http://www.evanmiller.org/sequential-ab-testing.html for Evan Miller suggesting a similar procedure to mine, based on past conversations that I have had with him. The difference between that one and mine is that I focus on "make the best decision available, even though it may be only little better than a coin flip" while he tries to only decide when you can decide with confidence. In past conversations we have agreed that each other's procedures are reasonable, but they achieve different goals.

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#129
post #110

Earlier quoted context omitted.

We are clearly not on the same page. Because I think that 1 and 3 are rather different things, and you don't. In particular 1 consists of exact statements about the likelihood of 7 births in a row being mmmmmmf. By contrast 3 consists of statements about what Bill and Lorena's childbearing plans would have been if something different had happened. Those are very different types of statement. There is no connection be…

Okay, my last try. For Bayes' theorem, we need a theory of how the data is produced given the parameter of interest. Bill and Lorena's plans certainly influence what data I observe: in scenario two, I can never observe the data BBBBGGG, but in scenario one, I can. My point is that your first category is not "exact statements about the likelihood of 7 births in a row being mmmmmmmf", it is "exact statements about the…

If you include into the priors information about the likelihood of the next child being born, you will indeed get different absolute probabilities. But you will not get different relative probabilities unless your available priors create a correlation between birth order and the likelihood of different genders for the next child if it comes. And therefore the probability of having the next birth cancels out of Bayes' formula and you wind up with the exact same conclusions from the observed data.

You certainly DON'T wind up with anything like the factor of 8 difference that frequentist techniques will see!

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#130
post #8

Is the p-value really not the probability of your results being due to chance? Is that not a perfectly valid definition of it? I suppose 'chance' is a little hand-wavy, but isn't a p-value just the probability of your data given that your hypothesis is false? Isn't that literally and precisely the probability that they occurred by chance?

Imagine I handed you a 20-sided die. I claim it says 7 on every side, but I might be lying. You roll a 7. What are the chances it actually has 7 on every side? You can't actually say unless you either (1) roll the die more times, or (2) assume something about the probability that I gave you an all-7's die to begin with. Doing (2) is useless, because that exactly the question we are trying to answer. For example, supp…

rolling the die more times just lowers the P value. You still can't make a definitive statement on the probability that it's an all-7's die without assuming something about the prior probability.

However, if you can bound the prior probability on the low end, you can make meaningful answers. Let's say you think there's at least a 1-in-a-billion chance that you have an all-7's die. After five rolls in a row of 7, there's at least a 0.3% chance that you were handed an all-7's die. After six rolls there's a 6% chance, and so on.

Usually this quantitative analysis isn't formally done, since the priors can always be debated, but rather a very small P value is demanded for very unlikely events.

Post reply on HN