Live data from Hacker News

Beautiful Probability

readthesequences.com

31–40 of 67 posts

Re: Beautiful Probability

#31

Earlier quoted context omitted.

If you see two people roll a d20 and get a 20, you get to say "wow, that was unlikely" to both of them, even if one of them privately admits they were going to quickly re-roll their die if they got below a 10. What matters is their actual behavior (identical in the example) not their intentions. The d6 vs d20 version is different because their behavior is different.

Let's imagine that we ran it as a simulation and we ran it a million times. The two people would have a different distribution of results. If you ignore the intention, you ignore reality as if that intention were not a part of it. Do you not notice that your inference is less accurate using this line of reasoning? Does that not suggest that it's simply wrong?

What do you mean by 'results'?

They would not have different distributions of results on their first die roll.

They would have different distributions of results on their reported die roll.

If I am looking at their first die roll, the fact that they would have different reported die rolls doesn't matter!

Re: Beautiful Probability

#32
post #27

Earlier quoted context omitted.

Let's imagine that we ran it as a simulation and we ran it a million times. The two people would have a different distribution of results. If you ignore the intention, you ignore reality as if that intention were not a part of it. Do you not notice that your inference is less accurate using this line of reasoning? Does that not suggest that it's simply wrong?

This is well put. Coincidentally in the example the results are the same , but they need not be. given repeated experiments with the same intentions one may expect different distributions. However, one could just move the argument up a level and manufacture a case of different intentions leading to the same distributions and then ask the same question.

> Coincidentally in the example the results are the same , but they need not be.

The questions is whether we should draw different conclusions when the results are the same. I don’t think that anyone has any issues with drawing different conclusions when the results are different!

Re: Beautiful Probability

#33
post #27

Earlier quoted context omitted.

Let's imagine that we ran it as a simulation and we ran it a million times. The two people would have a different distribution of results. If you ignore the intention, you ignore reality as if that intention were not a part of it. Do you not notice that your inference is less accurate using this line of reasoning? Does that not suggest that it's simply wrong?

This is well put. Coincidentally in the example the results are the same , but they need not be. given repeated experiments with the same intentions one may expect different distributions. However, one could just move the argument up a level and manufacture a case of different intentions leading to the same distributions and then ask the same question.

Imagine you have a machine that rolls a d20 and lies if the die comes up 1-19, and tells the truth on a 20. Should you trust this machine usually? No. But if you can _see that the die comes up 20_ then you should trust it. The fact that it sometimes might lie doesn't mean that you should distrust the machine if you can see that in this case it's telling the truth.

Re: Beautiful Probability

#34

Earlier quoted context omitted.

> the number of roots even for the same quadratic equation may depend on "private" thoughts such as whether complex roots are of interest You are confusing ambiguity in a problem statement due to human language being imprecise with two well-specified identical experimental results having different results due to the intentions of the human carrying them out. Is arithmetic a religion because there's "one true way" of…

I can think of at least two ways to add integers.. the categorical way that applies a mapping from the set into itself, and the set-theoretic way that deals with unwrapping and rewrapping successor relations. The latter is sometimes resorted to in heavily-relational contexts like Datalog.

Yes, this is addressed in the original article... there are multiple "lawful" ways of adding integers which all give the same results, and likewise in probability all "lawful" ways of analyzing data should give the same results. If you have two different ways of adding numbers which give different results, one is not lawful.

Re: Beautiful Probability

#36
post #7
post #5

So you know when you believe something and then you update your belief because you get some evidence? Yeah, and then you stack some beliefs on top of that. And then you discover the evidence wasn’t actually true. Remind me again what the normative Bayesian update looks like in that instance. Unfortunately it’s turtles all the way down.

> you discover the evidence wasn’t actually true Not really going to vouch for the normative Bayesian approach, but you might just consider this new (strong) evidence for applying an update.

The precise claim (I believe) is that the prior update which you had, made some assumptions about the correct way to phrase your perceptions.

That is, you say, for the update, "the probability that this trial came out with X successes given everything else that I take for granted, and also that the hypothesis is true" vs. "the probability that this trial came out with X successes given everything else that I take for granted, and also that the hypothesis is false." So you actually say in both cases the fragment, "this trial came out with X successes."

What happens if it didn't really? Well, the proper Bayesian approach is to state that you phrased this fragment wrong. You actually needed to qualify "the probability that I saw this trial come out with X successes given ...", and those probabilities might have been different than the trial actually coming out with X successes.

OK but what happens if that didn't really, either. Well, the proper Bayesian approach is to state that you phrased the fragment doubly wrong. You actually needed to qualify it as "the probability that I thought I saw this trial come out with X successes given...". So now you are properly guarded, like a good Bayesian, against the possibility that maybe you sneezed while you were reading the experiment results and even though you saw 51, it got scrambled in your head and you thought you saw 15.

OK but what happens if that didn't really, either either. You thought that you thought that you saw something, but actually you didn't think you saw anything, because you were in The Matrix or had dementia or any number of other things that mess with our perceptions of ourselves. So you, good Bayesian that you wish to be, needed to qualify this thing extra!

The idea is that Bayesianism is one of those "if all you have is a hammer you see everything as a nail" type of things. It's not that you can't see a screw as a really inefficient nail, that is totally one valid perspective on screwness. It's also not that the hammer doesn't have any valid uses. It does, it's very useful, but when you start trying to chase all of human rationality with it, you start to run into some really weird issues.

For instance, the proper Bayesian view of intuitions is that they are a form of evidence (because what else would they be), and that they are extremely reliable when they point to lawlike metaphysical statements (otherwise we have trouble with "1 + 1 = 2" and "reality is not self-contradictory" and other metaphysical laws that we take for granted) but correspondingly unreliable when, say, we intuit things other than metaphysical laws, such as the existence of a monster in the closet or a murderer hiding under the bed or that the only explanation for our missing (actually misplaced) laptop is that someone must have stolen it in the middle of the night." You need to do this to build up the "ground truth" that allows you to get to the vanilla epistemology stuff that you then take for granted like "okay we can run experiments to try to figure out stuff about the world, and those experiments say that the monster in the closet isn't actually there."

Re: Beautiful Probability

#37

As other commenters have pointed out any given introductory chapter in a book on Bayesian statistics, including Jaynes’, is better exposition than this. I found _Probability Theory: The Logic of Science_ very easy to follow and very well-written. I had a similar experience when I finally found a copy of Barbour’s _The End of Time_ and discovered, much to my chagrin, that it wasn’t nearly as mystical or complicated as…

Here's a link: http://www.med.mcgill.ca/epidemiology/hanley/bios601/Gaussia...

And if you want to read what he has to say on the optional stopping problem, you can scroll down to page 196 (166 in page numbers) to the heading "6.9.1 Digression on optional stopping"

I don't personally think Jaynes is much easier to read than Yudkowsky, but he's definitely more rigorous.

Re: Beautiful Probability

#38
From _Probability Theory: The Logic of Science_:

> Then the possibility seems open that, for different priors, different functions r(x1,..., xn) of the data may take on the role of sufficient statistics. This means that use of a particular prior may make certain particular aspects of the data irrelevant. Then a different prior may make different aspects of the data irrelevant. One who is not prepared for this may think that a contradiction or paradox has been found.

I think this explains one of the confusions many commenters have; for an experimenter who repeats observations until they reach their desired ratio r/(n-r), the ratio r/(n-r) is not a sufficient statistic! But when we have an experimenter who has a pre-registered n, then ratio r/(n-r) is a sufficient statistic. However, in either case,

> We did not include n in the conditioning statements in p(D|θ I) because, in the problem as defined, it is from the data D that we learn both n and r. But nothing prevents us from considering a different problem in which we decide in advance how many trials we shall make; then it is proper to add n to the prior information and write the sampling probability as p(D|nθ I). Or, we might decide in advance to continue the Bernoulli trials until we have achieved a certain number r of successes, or a certain log-odds u = log[r/(n − r)]; then it would be proper to write the sampling probability as p(D|rθ I) or p(D|uθ I), and so on. Does this matter for our conclusions about θ?

> In deductive logic (Boolean algebra) it is a triviality that AA = A; if you say: ‘A is true’ twice, this is logically no different from saying it once. This property is retained in probability theory as logic, since it was one of our basic desiderata that, in the context of a given problem, propositions with the same truth value are always assigned the same probability. In practice this means that there is no need to ensure that the different pieces of information given to the robot are independent; our formalism has automatically the property that redundant information is not counted twice.

Re: Beautiful Probability

#39

From _Probability Theory: The Logic of Science_: > Then the possibility seems open that, for different priors, different functions r(x1,..., xn) of the data may take on the role of sufficient statistics. This means that use of a particular prior may make certain particular aspects of the data irrelevant. Then a different prior may make different aspects of the data irrelevant. One who is not prepared for this may thi…

That seems a bit long winded since this situation is a direct result of Bayes' theorem. It seems to me equivalent to say:

Bayes' Theorem holds because it can be proven. Therefore, situations can be constructed where considering identical data without considering priors gives nonsense conclusions. For example if we happen to know as a prior that P(outcome of experiment is a certain ratio) = P(experiment is completed) then that must be considered when interpreting the results.

Re: Beautiful Probability

#40

Earlier quoted context omitted.

Unlikely in what probability space? We only see one version of reality so the probabilities that we assign to any outcome are based on a prior choice of probability space. That is why the researchers' intent matters.

Both events have the same probability of happening; 1/20. The fact that the researcher intended to do something in a reality that didn't happen isn't relevabnt.

If you want to know whether a drug is more effective than placebo, the answer to that question depends on both the data collected in a study and the initial study design. There’s a reason why it’s meaningless to say “that was unlikely” after somebody says they were born on January 1, or after getting a two-factor code that is the same number six times. There’s nothing special about those particular events except for the fact that we noticed them. Since we live in a single instance of the universe where they have already happened, they have probability 1. At the same time, on any given instance they have probability 1/365ish or 1/10000. The difference between these two interpretations of the probability is the same difference as having a good experimental design vs a flawed experimental design where you repeat the experiment until you get the results you want to see.
Post reply on HN