Live data from Hacker News

How Bayes’ Rule Emerged Triumphant from Two Centuries of Controversy

mcgrayne.com

71–80 of 83 posts

Re: How Bayes’ Rule Emerged Triumphant from Two Centuries of Controversy

#71
post #68
post #28

Earlier quoted context omitted.

Thanks, fixed. It understand that it is the limit, but the question I am asking is whether it is possible to always get heads, an infinite number of times. If this is possible, then it is not true that the relative frequency always approaches 0.5 as the number of repetitions goes to infinity. I understand that the probability of obtaining heads an infinite number of times is zero, but does that mean it is impossible?…

"whether it is possible to always get heads, an infinite number of times" No if we actually make the experiment, for we cannot toss coin infinitely many times. Yes, if we are just considering and measuring all possible infinite sequences. Measure of such mathematically possible sequence is zero, like measure of a point somewhere in a circle is zero. The point exists and is a possible result of choosing a point, but i…

If an infinite sequence of all heads is a possible outcome, then it is not true that the relative frequency converges to 0.5 in the limit. It is obviously true for many infinite sequences but if all heads is a possibility then it is not true in general.

Just to clarify, I am talking about the frequentist's claim that in the long run the relative frequency becomes the probability, i.e. P(x) = lim[n -> ∞] nx / n where n is the number of trails and nx is the number of trails that yielded x. So again, if an infinite sequence of heads is a possible outcome, then that would mean P(heads) = 1 while the frequentist asserts that it necessarily must become 0.5.

Re: How Bayes’ Rule Emerged Triumphant from Two Centuries of Controversy

#72
post #64
post #28

Earlier quoted context omitted.

Thanks, fixed. It understand that it is the limit, but the question I am asking is whether it is possible to always get heads, an infinite number of times. If this is possible, then it is not true that the relative frequency always approaches 0.5 as the number of repetitions goes to infinity. I understand that the probability of obtaining heads an infinite number of times is zero, but does that mean it is impossible?…

There are different ways to "approach 0.5 as the number of repetitions goes to infinity": https://en.wikipedia.org/wiki/Convergence_of_random_variable... I think in this case you have "convergence in probability" (but I've not read carefully the discussion). There is a stronger form, "almost sure convergence", where there is exact convergence with probability one (i.e. in some cases there is no convergence, but those…

That is exactly what I mean. We try to establish what a probability of 0.5 means, i.e. that the relative frequency converges to 0.5 if the number of trails goes to infinity, but that is not exactly true because there is a (possibly vanishingly small) set of outcomes where the relative frequency does not converge to 0.5. This in turn forces us to state that the relative frequency only converges with high probability, but now we have used a probability in our definition of probability, i.e. we made some kind of circular argument. I just skimmed the article you linked and I am not sure if I read it before so I will read it again, but from a first glance it does not look like any definition in there is suitable to avoid the problem.

Re: How Bayes’ Rule Emerged Triumphant from Two Centuries of Controversy

#73
post #37
post #28

Earlier quoted context omitted.

Thanks, fixed. It understand that it is the limit, but the question I am asking is whether it is possible to always get heads, an infinite number of times. If this is possible, then it is not true that the relative frequency always approaches 0.5 as the number of repetitions goes to infinity. I understand that the probability of obtaining heads an infinite number of times is zero, but does that mean it is impossible?…

> the question I am asking is whether it is possible to always get heads, an infinite number of times No. By my definition (lay, may be wrong), the probability is the ratio of heads to tails to which we converge at the limit (infinity). The only way for it to converge to 0 is if the probability of heads was 0 to begin with, which would contradict the initial assertion that it was a fair coin. By the way, from your pr…

No. By my definition (lay, may be wrong), the probability is the ratio of heads to tails to which we converge at the limit (infinity). The only way for it to converge to 0 is if the probability of heads was 0 to begin with, which would contradict the initial assertion that it was a fair coin.

This seems problematic to me. You can not perform an infinite number of coin tosses and know to what the relative frequency converges. So how do you conclude that a fair coin has probability of 0.5 for heads then? I can't really put my finger on it, but that argument is somehow circular. The probability is what the relative frequency converges to and I can not get anything other than 0.5 because that means to probability would have to have been not 0.5 to begin with. I can really only say that I disagree, I just can't pin it down exactly.

Are you familiar with the argument of whether 0.999... equals 1? (Spoiler: it does) Some people have a really hard time coming to terms with it, and the fundamental problem for many comes down to the same difficulty of reasoning about infinity.

I am aware of that and it seems totally obvious to me. And I also certainly know that infinity is not just a really huge number, but I am not sure I internalized that well enough to not make any mistakes because of the difference. Actually I am pretty sure I make mistakes because of that.

Re: How Bayes’ Rule Emerged Triumphant from Two Centuries of Controversy

#74
post #26

I could never quite understand the divide between Bayesian statistics and frequentist statistics. Both seem to be ultimately about counting the frequency by which something occurs and normalizing this frequency with respect to the number of all possible outcomes. Bayesian statistics essentially is concerned with the application of the Bayesian updating technique by which one can iteratively improve a distribution ove…

Bayesian and frequentist approaches ultimately have a different notion of probability. In the frequentist approach, a probability of 10% means that if you repeat an experiment many times, roughly 1 out of 10 times you will observe an event. In Baysian statistics, a probability of 10% means that you are that certain about the event happening. So you would be willing to bet at 10 to 1 odds on the event happening. There…

Another way to look at it is this way:

P(H|D) = P(D|H) P(H) / P(D)

Bayesians are interested in the probablity of various hypotheses h in H given some data D.

Frequentists calculate the probability of some data given a hypothesis (p-value is not strictly a probability but it can be one - it is ALWAYS a measure of extremity of data coming from the assumed hypothesis, which can be considered a relative probability).

Most interesting to me is that the Bayesian formula includes P(D|H) which is basically what the frequentists are calculating. In this sense, the question Bayesians answer is far closer to what we want to ask and far more powerful. In practice, the frequentist approach is often more than enough, though. The tradeoff is computability and simplicity.

Re: How Bayes’ Rule Emerged Triumphant from Two Centuries of Controversy

#75
post #72
post #64

Earlier quoted context omitted.

There are different ways to "approach 0.5 as the number of repetitions goes to infinity": https://en.wikipedia.org/wiki/Convergence_of_random_variable... I think in this case you have "convergence in probability" (but I've not read carefully the discussion). There is a stronger form, "almost sure convergence", where there is exact convergence with probability one (i.e. in some cases there is no convergence, but those…

That is exactly what I mean. We try to establish what a probability of 0.5 means, i.e. that the relative frequency converges to 0.5 if the number of trails goes to infinity, but that is not exactly true because there is a (possibly vanishingly small) set of outcomes where the relative frequency does not converge to 0.5. This in turn forces us to state that the relative frequency only converges with high probability,…

There is no circularity. You're discussing the real-world interpretation of the probability of an event in terms of long-term frequencies. The probability in the definition of convergence is a well-defined mathematical concept that has nothing to do with frequencies.

Edit: thinking more about it, I agree that the passage from mathematical probability to real-world probability is not very satisfactory. But it's not really surprising, the very notion of "real-world" is troublesome.

Re: How Bayes’ Rule Emerged Triumphant from Two Centuries of Controversy

#76
post #67

Earlier quoted context omitted.

What would you respond to my other comment downthread? https://news.ycombinator.com/item?id=11985863

If you are able to think of the different realities that would lead to us having this discussion today, some of them with life being originated on planet Earth and some of them with life coming from elsewhere, and you're able to reason about the relative frequency of these two kinds of realities, you definitely have more imagination than me. And I don't know what do you gain with that. Does your frequentist interpret…

It is already acknowledged that the initial guess can be wrong. We don't need to consider all possibilities to come up with a prior distribution because that is basically the 'trick' of the Bayesian updating scheme: Just start somewhere and we'll get closer to the true distribution by collecting more data and by using it the most logical way, namely by solving for the posterior. No imagination needed. It is more efficient to distribute the initial probability mass according to our best guess using a lot of imagination, but that itself is basically Bayesian updating because human reasoning is approximately Bayesian (or Bayesian with a noisy prior/bias).

A physical interpretation might be that all the other realities are realized in terms of the Many-worlds interpretation or in terms of Tegmark's level 4 multiverse. Without much information we cannot really nail down which reality we find ourselves in (cf. the sleeping beauty problem and Boltzmann brains). But we can use Bayesian updating to become more certain about what reality is about.

What do I gain from that? I am not entirely sure, but I find a probability is just better interpretable as frequency or fraction compared to a subjective quantity.

Re: How Bayes’ Rule Emerged Triumphant from Two Centuries of Controversy

#77
post #74
post #26

Earlier quoted context omitted.

Bayesian and frequentist approaches ultimately have a different notion of probability. In the frequentist approach, a probability of 10% means that if you repeat an experiment many times, roughly 1 out of 10 times you will observe an event. In Baysian statistics, a probability of 10% means that you are that certain about the event happening. So you would be willing to bet at 10 to 1 odds on the event happening. There…

Another way to look at it is this way: P(H|D) = P(D|H) P(H) / P(D) Bayesians are interested in the probablity of various hypotheses h in H given some data D. Frequentists calculate the probability of some data given a hypothesis (p-value is not strictly a probability but it can be one - it is ALWAYS a measure of extremity of data coming from the assumed hypothesis, which can be considered a relative probability). Mos…

Interesting, I've never thought about the likelihood as a confidence, but it makes sense. But sometimes the confidence also seems to reflect the opposite of extremity (e.g. for the null hypothesis).

Re: How Bayes’ Rule Emerged Triumphant from Two Centuries of Controversy

#78
post #15

Earlier quoted context omitted.

Did you find out the base statistics for fatalities in skydiving? That that particular business had fatalities may just indicate that they do a lot more volume or more advanced types of skydiving than other companies (in which case they may actually be safer than the norm). For me, one of the big differences Bayes rule makes is spotting that information nearly always needs context to understand it correctly. You don'…

Yeah I thought about this and could find a news article that said they did around 5000 jumps a year since they started in 2001 (and they're located in Las Vegas so they I assume most of their customers are first timers). So that is 75000 jumps. In the news I could find two fatalities, and two serious injuries (but let's ignore those for now). Apparently, statisticians use the term "micromorts" which means a death per…

You aren't applying Bayes correctly. As a general guide, you need to add the word "given" to your problem statement, assign two "situation A" and "situation B" variables, and then work the math.

For example, assigning variables:

A = You died B = You're skydiving at that particular dropzone

Then: "What is the risk of mortality GIVEN that I am skydiving at this drop zone?" (P(A|B))

"What's the chance that I'm skydiving at this dropzone, given the fact that I died?" (P(B|A))

"What's the risk of my mortality while skydiving?" (P(A))

"What is my probability of skydiving at this dropzone?" (P(B))

P(A | B) = (P(B | A) * P(A)) / P(B)

So you calculated P(A | B) in a non-Bayesian way, without finding out the other information to calculate it using Bayes, and then stopped there, like most people do. This is why Bayes is often difficult for people to understand and apply correctly, and, honestly, it's probably not the equation you want for the situation you're looking at.

Another approach -- and this seems to be the one you want -- is to calculate a 95% confidence interval using a binomial distribution, to find out if their statistics are really anomalous. Death is a relatively rare event, and, even if they're distributed perfectly randomly, you'll find odd-looking clusters here and there.

To figure out if it's anomalous, many people would use the normal approximation to the binomial confidence interval, which would be wrong -- the probabilities of death are so relatively tiny, that they can't be approximated normally (rule of thumb is P(A) x P(not A)x sample size > 5 to use the normal distribution, which this fails), so we need to do an exact calculation. I've done this by hand before, but it's a pain in the butt. That's why we have calculators!

http://epitools.ausvet.com.au/content.php?page=CIProportion (If you don't trust it, you can use another one)

You enter your numbers:

sample size = 75000

number of deaths = 2

This gives the exact binomial confidence interval as: [3.23e-06, 9.633e-05]

This means that their actual death rate could be anything from .00000323 (that's 3 deaths per million) to .000096 (that's nearly 100 deaths per million)

Clearly, you cannot say with any meaningful level of confidence whether or not this drop zone is safer, or less safe, than average. Sorry!

Edit: In my Bayes example earlier, I realized that you could actually use it in an interesting-ish way (I guess?) to find out an unknown: "What are the odds that I was skydiving at this particular dropzone, given that I died skydiving?"

So: A = You're skydiving at that particular dropzone

B = You died :(

P(A | B) = (P(B | A) * P(A)) / P(B)

We know that: P(B | A) = 2/75000 = .0000267

P(A) = their skydives / all skydives = (75000 / 3,300,000 x 15) = 0.0015

Note: I found out that there were 3.3 million skydives in the US in 2012, so, let's extrapolate that out 15 years as a rough approximation, and limit us to just the US.

P(B) = .000009

Then:

P(A | B) = (.0000267 * .0015) / .000009 = 0.00445

So, if you died, then the probability that you were skydiving at that dropzone is 0.4%!

Note though, that there's some uncertainty in the calculation of P(B | A) (the probability of dying while skydiving at that dropzone) which you need to use confidence intervals, above, to actually figure out. Anyway, I certainly clarified some of my thoughts while writing this, and I hope that it helps you too, figuring out when to use confidence intervals, and when to use Bayes, and why it matters!

Re: How Bayes’ Rule Emerged Triumphant from Two Centuries of Controversy

#79

Brexit is the best example so far. That painful dissonance between so called reality and these probabilistic models.

Probabilities makes sense only with absolutely certain things like a fair coin or a dice. In cases where there is no absolute certainty about how many sides or dimensions your "dice" has and that it is not biased and that there is no other forces or factors in play probability ceases to make sense. Probability of A, given B becomes meaningless when either A or B aren't precisely defined (like in the case of a "fair c…

That's wrong. With the Bayesian interpretation of probability, you can assign probability to any event. In fact all certainty about belief is just probability in disguise.

There were betting markets for the brexit vote. They assigned 25% probability to brexit. Of course markets aren't perfect. But anyone who really believes they know better should be able to get rich off them. And somehow that doesn't happen. So they are the best estimates of probability we have.

Re: How Bayes’ Rule Emerged Triumphant from Two Centuries of Controversy

#80
post #15

Earlier quoted context omitted.

Yeah I thought about this and could find a news article that said they did around 5000 jumps a year since they started in 2001 (and they're located in Las Vegas so they I assume most of their customers are first timers). So that is 75000 jumps. In the news I could find two fatalities, and two serious injuries (but let's ignore those for now). Apparently, statisticians use the term "micromorts" which means a death per…

You aren't applying Bayes correctly. As a general guide, you need to add the word "given" to your problem statement, assign two "situation A" and "situation B" variables, and then work the math. For example, assigning variables: A = You died B = You're skydiving at that particular dropzone Then: "What is the risk of mortality GIVEN that I am skydiving at this drop zone?" (P(A|B)) "What's the chance that I'm skydiving…

This is great. Thank you
Post reply on HN