Live data from Hacker News

Gambler’s Fallacy and the Regression to the Mean

theness.com

71–80 of 96 posts

Re: Gambler’s Fallacy and the Regression to the Mean

#71

Earlier quoted context omitted.

On the gripping hand, coming up red 10 times in a row is evidence that the odds of red and black aren't even.

Strangely, ten reds in a row is not particularly good evidence of an unfair wheel. For an American wheel, the odds of red in a single spin is 18/38= 0.4737. So ten in a row is 0.4737^10, which is about 0.0005689. So one should expect five or so such sequences in ten thousand spins. Or one in every two thousand spins. If a casino spins its roulette wheel 200 times a day, you would expect to see a sequence of ten reds…

Depends on the situation. At a well-regulated, long-lived casino with lots of tables, it's not good evidence.

But if you see it run by carnies at the county fair where the wheel will be gone tomorrow?

Re: Gambler’s Fallacy and the Regression to the Mean

#72
post #30

> That’s a great question, and the answer is a definite no – they are not in conflict. Again, the pressure to think that the past influences future independent events is powerful. Regression to the mean is not a power in the universe that ensures that statistics work out in the end, it is purely a probability. I find TFA's argument about the gambler fallacy not being associated with regression to the mean quite hand…

But "streaks" are irrelevant. Reversion to the mean doesn't tell you things across trials, it tells you things about individual trials. Try this on: What is the "mean" red-black on the roulette wheel? There isn't one. The only way you can reasonably expect the past streak to have an impact on the next spin, is if you've concluded the streaks are sufficiently unlikely to cause you to judge the wheel not to be fair. If…

>If it's random, then you have no recourse but to post-hoc description. There is -no- predictive weight.

Well, a "post-hoc description" of empirical observations is how we make predictions. We gather cases (and their probabilities) post hoc, and predict future results based on that.

Why shouldn't it be applied in this case, especially when here not only we can calculate the probabilities post-hoc, but even use probability to pre-calculate those post-hoc probabilities for every series of streaks?

>The only way you can reasonably expect the past streak to have an impact on the next spin

Regression to the mean is a reason to expect past streaks to have an impact on future spins. Doesn't have to be some force that specifically affects the individual next spin, but it's a probability that affects (and moves toward the mean) the outcome of next spins.

Intuitively speaking, why shouldn't the gambler use it?

If they shouldn't, "because spins are independent" seems too hollow as an answer, because we have independent spins but also a statistical property for their aggregation.

I guess we could test it empirically: in a fair sequence of coin tosses, would a gambler always betting on tails after seeing N (for e.g. N > 4) consequtive heads have an advantage, over one always betting randomly on the same N+1 toss?

But regardless of that, my point is that intuitively is quite justified, and the explanations why it's not seem more hand-wavy.

>Try this on: What is the "mean" red-black on the roulette wheel? There isn't one.

Well, (ignoring non red-black case), there's not a mean value, but for the distribution there's an expected long term tendency of a balanced red and black outcomes.

>... And editing your comment to refer to groups of throws is disingenuous. It would have been more honest to modify your case in a reply.

I didn't modify my case, I expressed it more clearly. What exactly was disingenuous? I more often than not post a quick comment (HN habbit, as sometimes old edit forms expire), and then re-read it and quickly improve it as I re-read it, find typos, better wording, etc. Didn't even think anybody would read it and answer in the time it took me to refine it.

And, how about it, you can always answer to my final, improved, formulation. Does it have any historical interest if you answer to a worse worded case?

Re: Gambler’s Fallacy and the Regression to the Mean

#73
post #22
post #14

Earlier quoted context omitted.

"which to me makes intuitive sense" Your intuition is either very good or complete bollocks, or at least worryingly odd 8) Monty Hall is a really clever problem and worth studying in some depth. Whenever I've encountered it, the rules are always given without ambiguity. Even so, it is very hard to get to the bottom of the probabilities. You can reason your way through it and possibly get to the right answer, unaided.…

Maybe I'm just misunderstanding something then? I'm not trying to be dismissive or act like I think I have some special intuition here. It really does just seem straightforward. Imagine the problem this way: The host of a game show presents you with 100 doors, behind one of which is a prize. You pick one, and you know that your odds of having chosen the correct door are 1 in 100. The host, who knows where the prize i…

What you're missing here is that you already have enough intuition to translate it into a more obvious statistical problem, but most people come across this as blank slates, and under very different framing.

Even when relayed the right way and the wording isn't necessarily unclear or misleading per se, it is still framed in a particular way and context that pushes you to believe your choice is inconsequential, less about probability and math, and more for psychological games and dramatic effect in the context of a TV show.

In your example, the fact that the presenter opens all 98 doors and not just one is telling. If they only opened 1, and then told you to stick to your original or open one of the remaining 98 doors, that he picks for you, would you change?

This is what you're missing, I think. Most people don't even get to the point where they reframe the "would you change" question to "what are the chances my original choice was correct vs not". Most people get stuck on "He knows something; he clearly demonstrated very theatrically that he knew the car wouldn't be behind door No.2. And now he's asking me to swap my choice for door No.67. Why 67? Is he trying to trick me? Is this a psychological trick to get me to forfeit my door, which may be the right one? Should I trust him? After all, HOW IS MY DOOR ANY DIFFERENT THAN ALL THE OTHERS?"

(the fallacious thinking, obviously, lying in the last part here)

Re: Gambler’s Fallacy and the Regression to the Mean

#74
post #6

So the author presents the Monty Hall problem this way (very explicitly saying that the host knows where the prize is an will not reveal it): > You are given a choice of three doors, behind one is a prize. You can choose one door. The host of this game, who knows where the prize is, then opens one door without a prize (again – they know where the prize is and deliberately choose one of the unchosen doors without a pr…

> If the host is choosing a door randomly and it doesn't happen to contain the prize, your odds don't improve if you switch your answer.

No.

If the host chooses a door randomly and it doesn’t happen to contain the prize, the outcome is exactly the same as if the host does know where the prize is. Look up “conditional probability” for how to calculate this. The host’s knowledge has no effect other than to prevent the awkward outcome in which the host opens the door with the prize.

Even if the host did open the door with the prize on occasion, the outcomes would still be the same if the rule was that the whole game starts over if the host accidentally reveals the prize. (Well, the game would take longer to finish, but that has no effect on the probability of winning.)

Re: Gambler’s Fallacy and the Regression to the Mean

#75
post #30

> That’s a great question, and the answer is a definite no – they are not in conflict. Again, the pressure to think that the past influences future independent events is powerful. Regression to the mean is not a power in the universe that ensures that statistics work out in the end, it is purely a probability. I find TFA's argument about the gambler fallacy not being associated with regression to the mean quite hand…

After successive black streaks about the only thing you can say is maybe the wheel is biased toward black. Assuming no green, and a perfectly unbiased wheel, black if a fifty fifty chance. Which means the next infinity sounds half will be black. If there were three blacks in a row, it doesn't mean red is more likely, as the rest of the spins half will be black.

>Assuming no green, and a perfectly unbiased wheel, black if a fifty fifty chance. Which means the next infinity sounds half will be black. If there were three blacks in a row, it doesn't mean red is more likely, as the rest of the spins half will be black

Well, the point is:

"black if a fifty fifty chance" and "the next infinity sounds half will be black"

doesn't sit well with:

"if there were three blacks in a row, it doesn't mean red is more likely".

The roll is independent, but in some sense either the roulette is unfair or red will increasingly be more likely.

In other words, since there is a statistical property working across many/infinite rolls (regression to the mean, black/red equally likely), this property makes sense to also somehow affect the possibility of individual rolls (as individual rolls are still rolls within an aggregate).

Re: Gambler’s Fallacy and the Regression to the Mean

#76
post #5

> “I know that the fact that the roulette wheel has come up red 10 times in a row tells me NOTHING about spin #11. On the other hand, I know that over time, there will be just as many black spins as red spins, so at least intuitively, a black spin seems at least a little more likely to come up next in order to push that ratio back towards 50/50. Are these two principles actually in tension with each other? If not, ho…

I find it interesting that saying "a black spin is needed to push the ratio back towards 50/50" feels like a sensible thing to say. What role do context and framing play?

Consider the first one hundred spins breaking 60 red, 40 black. That is a little odd, but not outrageously so. The ratio is 60/40 = 1.5

Suppose that the next one hundred spins break 50/50. There are no surplus blacks to push the ratio towards 50/50. So what is your intuition? It seems clear enough that the totals are 110 red, 90 black, for a ratio 110/90 = 1.22...

Having imagined the next hundred spins vividly, we are comfortable saying "Unless the surplus reds keep coming, the ratio is going to drift back towards 50/50 all by itself.". There is a striking clash between different "things we might say" seeming sensible, depending on the vividness and detail with which we imagine the situation.

Re: Gambler’s Fallacy and the Regression to the Mean

#77
post #29

Earlier quoted context omitted.

Your intuition is wrong and it doesn't matter what the host knows before showing you the door that doesn't have the prize. Here's some explanation: https://betterexplained.com/articles/understanding-the-monty...

How can they reliably show you the door that doesn't have the prize, if they don't know where the prize is? They can't, which means the host must know, which means it must matter what the host knows?

[deleted]

Re: Gambler’s Fallacy and the Regression to the Mean

#78
post #47
post #30

> That’s a great question, and the answer is a definite no – they are not in conflict. Again, the pressure to think that the past influences future independent events is powerful. Regression to the mean is not a power in the universe that ensures that statistics work out in the end, it is purely a probability. I find TFA's argument about the gambler fallacy not being associated with regression to the mean quite hand…

HH is less likely than H (.25 vs .5), but HT is also less likely than H and by the same amount. HH and HT are equally likely. This extends to HHHHHH and HHHHHT, and so on. There’s no way to frame the gambler’s fallacy that makes it a better bet. The probability that there will be one T somewhere in the sequence is a lot higher than the probability that there will be no T. That's only because there are a lot of somewh…

>HH is less likely than H (.25 vs .5), but HT is also less likely than H and by the same amount. HH and HT are equally likely. This extends to HHHHHH and HHHHHT, and so on.

Well, kind of. There are 2^6 patterns with 6 tosses, and only a few of them have some recognizable pattern (are compressible). HHHHHH is the worst of them. It might be "equally likely" to any other 6-toss pattern seen from a fair coin, but it's also a great example to make one suspect a totally rigged (same sided H) coin.

Now, on average a fair coin would have roughly balanced H and T. So the (much more numerous and thus probable) series where H and T are mixed (e.g. HTTHHT..., HTHTHH..., etc...), are more probable than a large streak of H.

>The probability that there will be one T somewhere in the sequence is a lot higher than the probability that there will be no T. That's only because there are a lot of somewheres for the T to be, but only one way for there to be no T.

A, my sentiments exactly!

>But once you have already seen 99 H, you still have precisely zero information about the next fair flip, and there's no way around it.

That's my argument: you don't have zero information.

You know that the next fair flip is random - sure.

But you also have the information that large series of tosses end up balanced in the long run, which is a kind of push (given fairness) towards making a T after tons of Hs increasingly likely.

Why should this be discarded?

Re: Gambler’s Fallacy and the Regression to the Mean

#79
post #78
post #47

Earlier quoted context omitted.

HH is less likely than H (.25 vs .5), but HT is also less likely than H and by the same amount. HH and HT are equally likely. This extends to HHHHHH and HHHHHT, and so on. There’s no way to frame the gambler’s fallacy that makes it a better bet. The probability that there will be one T somewhere in the sequence is a lot higher than the probability that there will be no T. That's only because there are a lot of somewh…

> HH is less likely than H (.25 vs .5), but HT is also less likely than H and by the same amount. HH and HT are equally likely. This extends to HHHHHH and HHHHHT, and so on. Well, kind of. There are 2^6 patterns with 6 tosses, and only a few of them have some recognizable pattern (are compressible). HHHHHH is the worst of them. It might be "equally likely" to any other 6-toss pattern seen from a fair coin, but it's a…

> Why should this be discarded?

Well, it's a restatement of the Gambler's Fallacy, but I'll taboo using that as an argument because it begs the question (but for reals!)

I thought about this for a while, and here's the best I can do for right now. It's like winning the MegaMillions lottery on Saturday, and then winning $1 on a scratch ticket on Sunday. You'd go running around screaming on Saturday, but you wouldn't on Sunday. (Would you?) Because the impossible already happened. Hitting $1 on a scratch-off is never remarkable (unless you're 8), including on the day after you won the MegaMillions. It's just a somewhat anti-climactic coincidence.

Re: Gambler’s Fallacy and the Regression to the Mean

#80
post #64

Earlier quoted context omitted.

Think of it slightly differently: Imagine you pick a door, and then the host, who has no idea what's behind the doors, also picks one. Then the remaining door opens and reveals that there's nothing behind it. Yes, if the host were to pick a door, not tell you which it was or reveal anything else and then offer to switch your choice for their choice, there would be no difference in the odds of each choice. That just h…

I disagree with your conclusion that the odds remain 2/3 even if the host doesn’t know where the prize is. If that were the case, it would mean that if the same scenario is repeated many many times, then in 2/3 of the cases you would still pick the door with the prize with the same strategy. But what about the instances where the host picks the door with prize before you even get a chance to pick the other door? Thes…

Do you still have the opportunity to switch if he picks the winning door?

Even if you could, you already know that you're losing, unless he doesn't show you what's behind it.

If the game is still going, you know that he picked an empty door regardless of whether he knew he was doing it

Post reply on HN