Live data from Hacker News

Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

statmodeling.stat.columbia.edu

131–140 of 243 posts

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#131

The thing about this election is that what if Trump himself is some sort of wildcard that can't really be properly forecasted in the polls? Why is there so much fascination with polls to begin with? I understand that there are betting markets, but it seems sort of silly. If you had a 100% accurate poll, for instance, then what would be the purpose of the actual election?

Now that you bring that up, if we had a 100% accurate poll that would be really good for productivity wouldn't it? Perhaps it wouldn't give voters the same feeling of self-determination but it'd save a lot of resources in fundraising, going out to vote, counting votes

I mean technically the election is a 100% accurate poll. You get data from every voter.

To run a 100% accurate poll would require you to sample every voter, so it would literally be an election.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#132

Every single one of these models will break down this year. We are living in an unprecedented time. I can't understand how we can model how many people will vote, when we don't even know how many people have moved out of cities this year. Half of my friends have left San Francisco - if as many people left Philadelphia, The Twin cities, Milwaukee or Pittsburgh, then that really effects the outcome.

I've made similar points on topics like this, and bar none, every single time, it is downvoted into oblivion. People seem to have a difficult time with forecasting data that goes against their preferred outcomes.

If you'd like to not be downvoted into oblivion you'd do well to provide an alternate theory. What's your prediction for the upcoming election? Why? What's your methodology? Or is your point that forecasting is futile? I can't tell. There's too much emotion directed at 538/etc. for me to suss out what your point is besides disliking 538/the media/etc, which just isn't particularly helpful for the discussion.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#133

The first problem with the model is it grossly missed the last election. I still think it’s good reading.

Are you referring to the 2016 election? If so, you are wrong. 538 gave Trump a higher chance of winning than pretty much every independent pollster.

So they were "very very" wrong rather than "very very very" wrong like other pollsters? I think you proved the OP right rather than wrong.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#134
post #7

See also this Twitter thread from Nate Cohn (who has no dog in the fight): https://twitter.com/Nate_Cohn/status/1320042092694065153

I think the third post in that thread really nails the underlying issue here. We, as humans, know that it would be silly to think that some configurations could happen. We know there's no reason to expect Biden to win Alabama outside of him winning nearly every state.

A statistical model only has a vague idea of context/the real world. It looks at polls (and probably not really that many polls of Alabama or Mississippi or Alaska) and sees that, statistically, Biden should win 3% of the time or so.

It doesn't have a specific world set of events in mind that would cause that, it just knows that that's how the numbers go, and thus may lead to weird circumstances in the grander results because it has to make the world match the numbers in these small corners.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#135
post #60

Earlier quoted context omitted.

> I'm also deeply skeptical of the notion that polling is even remotely correlated to actual results. But it has been strongly correlated to the results in basically all elections so far in all democracies on the planet. Taking 2016 as an example there has been a very strong correlation between polling and the results. The national polling averages were only 3 points off from the actual result. If that's not correlat…

Polarization and the unacceptability of publicly saying "I voted for X" also didn't really exist prior to 2016. The fear of getting doxxed, combined with a record low level of trust in institutions and the media, leads to skepticism toward answering polls truthfully, IMO.

While polarization is bad now, it's nowhere close to historical extremes, and it's not even as bad as it was during the post-Vietnam era only a few decades ago.

As for fear of doxing: plenty of people openly supported and voted for George Wallace (a noted white Supremacist) back in the day, and even Roy Moore (accused pedophile) just 2 years ago. Proud Boys members openly pose for the cameras even as they espouse racist views, and QAnon members brag about being part of QAnon.

The purported shy Trump voter? Doesn't exist. https://fivethirtyeight.com/features/trump-supporters-arent-...

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#136

Earlier quoted context omitted.

I've made similar points on topics like this, and bar none, every single time, it is downvoted into oblivion. People seem to have a difficult time with forecasting data that goes against their preferred outcomes.

If you'd like to not be downvoted into oblivion you'd do well to provide an alternate theory. What's your prediction for the upcoming election? Why? What's your methodology? Or is your point that forecasting is futile? I can't tell. There's too much emotion directed at 538/etc. for me to suss out what your point is besides disliking 538/the media/etc, which just isn't particularly helpful for the discussion.

I wish it were that simple. HN just seems to have gone the way of reddit, where downvoting is disagreement. And of course, 3 minutes after I post this comment, it's downvoted.

I don't have a methodology, because I'm not a pollster with dozens of people at my disposal. I am just bemused and annoyed that things like 538 continue to be taken seriously when they continue to ignore sociological, historical, and cultural factors in favor of an overly-complex quantitative model.

Re: the upcoming election. I don't think we can be sure, yet. Certainly it will be close, and the Biden at 90% to win estimations make no sense to me. Biden is a much weaker candidate than Hillary and he continues to make blunders (i.e. I guarantee that his comments on fracking in the last debate just lost him Pennsylvania.) Trump seems to be finding a lot of allies in strange places, e.g. African-American celebrities. That may be an isolated incident, or it may signal some big unexpected changes.

At this point my estimation is Trump-Biden 55-45, for the simple reason that people tend to vote for economic issues and Trump has a better "perception" on this issue. "It's the economy, stupid." as James Carville put it.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#137

Earlier quoted context omitted.

> That's why this time around the pollsters made sure to be more thorough in their polling. "This time is different." I've heard that enough times to be highly skeptical. I'm also deeply skeptical of the notion that polling is even remotely correlated to actual results. Cultural and historical trends play a drastically higher role and are almost always left out.

The shy Tory factor[0] is likely to be even stronger this year than in 2016 when it comes to polls. After 4+ years of being harangued and called every name in the book by the vast majority of national culture (movies, music, TV, news, social media, news-entertainment), along with increasingly hostile projects such as https://donaldtrump.watch/ I imagine less of his more subdued supporters are going to be honest with…

Anecdotal, but I have never met a Trump supporter that was shy about who they were voting for, this year or in 2016.

If anything, Trump supporters have been extremely vocal about who they were supporting, to the extent that they frequently violate social norms and try to takeover events and gatherings to make their political affiliations known, like this week with Among Us.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#138
One thing to note: 538 uses a t-distribution (and calls out on regularly in their podcast that this yields much heavier tails than a normal distribution). Even 40,000 samples is not enough to characterize the tails.

So, it seems to me that the entire article is predicated on a faulty conjecture, namely that 538 uses a mixture of a normal distribution with an independent heavy-tailed one. (It's not explicitly stated what the author thinks the base model is, but I think "normal" is a reasonable guess.)

I'd be interested in seeing a reverse-engineering analysis of 538's choice of distribution parameters, and extrapolation from there to see if these pathologies still arise with (much) larger samples.

...

That said, ultimately, the choice of how fat to make the tails is a modeling decision, and how the models behave outside the regime of interest isn't as important as how they behave within the operating region. There are key ways we can evaluate goodness of fit once we have results (e.g. bias, MSE) which we can use to determine just how wrong the model was as a predictor, and chances are pretty good that we won't see, say, Trump winning NJ, so we won't actually be able to validate the tail correlation with the vote in PA. But we will be able to validate the correlation in margin between PA and NJ.

Maybe 538's tails are too fat, and every prediction in the 80-95% range ends up going as predicted. Or maybe they're not fat enough, and some races in the 99% bucket end up going the opposite way. Point is, we won't know for sure which models were the best predictors until we can verify the predictions.

(see: all models are wrong, etc. Newtonian mechanics work great as long as your objects are big and slow, for instance.)

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#139

Earlier quoted context omitted.

That's an interesting thought, but San Francisco aside, it seems like most people moving out of cities are just moving to the suburbs of those cities, which shouldn't affect the presidential election calculus?

I know people that have moved, but forgot to register. Its a fact of life that voter registration will take the back seat when managing more complicated things in life - like a move

I don't think the type of person that forgets to register is the type to reliably show up at the polls, so I doubt this is going to have a big effect.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#140
post #19
post #12

Earlier quoted context omitted.

Yeah, but what if we could never roll that particular die again? I think that's what he's talking about.

You treat the forecaster as the dice, not a particular forecast from them. You can measure their forecasts with reality.

Yes, I'm not defending the original guy, I'm just stating what I believe his argument to be.

He's basically saying he doesn't believe the die is fair/unweighted.

So stating the odds of a fair die is kind of immaterial to his point. We need to demonstrate to the guy that the die is fair.

Someone else posted a link (I think) of 538 going over on how accurate they were. Whether their odds bore out. Here it is: https://projects.fivethirtyeight.com/checking-our-work/

Basically what they did was bucket every prediction by odds. If they predicted 70/30, it went in that bucket. And they're "right" at about the rate of their predictions. In other words, for every 70/30 prediction they made, the people/teams with 70% chance to win, won about 70% of the time.

That shows that 538 in this case is pretty decent at calculating odds.

Post reply on HN