Live data from Hacker News

Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

statmodeling.stat.columbia.edu

101–110 of 243 posts

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#102
This is undoubtedly useful to know when playing around with different scenarios.

However, for my use of 538, I’m perfectly happy to ignore such scenarios (such as Trump taking New Jersey). I can call the election in his favour by myself in these scenarios without needing the model.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#103
post #77

Earlier quoted context omitted.

"He won" is not sufficient evidence upon which to do that.

Nowhere did I say it is. I simply think that, if one were following the right information, his win was not as unexpected as the coastal media presented it as being. Personally I would have put it about 60-40 Hillary-Trump.

You started off by calling the 2016 prediction a debacle, but now you're saying you would have put the odds about 12 percentage points differently. That doesn't seem like a big enough disagreement to warrant the kind of vehement criticism you're throwing around.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#104

Earlier quoted context omitted.

That's an interesting thought, but San Francisco aside, it seems like most people moving out of cities are just moving to the suburbs of those cities, which shouldn't affect the presidential election calculus?

I know people that have moved, but forgot to register. Its a fact of life that voter registration will take the back seat when managing more complicated things in life - like a move

I wonder if that is true for people who were also likely to vote? I know registering to vote was one of the first things I do whenever I move, and I've also voted in every election since I've been allowed to.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#105
post #17

Meh. If you fit a model and don't explicitly constrain against "un-physical" results like negative correlations, you'll end up with them. Constraining against them won't improve your models fit (usually by definition), and it doesn't always improve robustness (at least for situations near average)-- because they're acting to debias the model in ways that you otherwise don't have enough degrees of freedom to address.…

> If you fit a model and don't explicitly constrain against "un-physical" results like negative correlations, you'll end up with them.

The Economist model does exactly that, and all of their correlations are positive.

I recommend reading their methodology, they know what they're doing (I wouldn't say the same about 538). Andrew Gelman has developed some of the Bayesian methods and software that people like Nate Silver use, he's the main author of what's considered a reference book on Bayesian statistics.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#106
post #52

Somewhat interesting, however the guy lost me more and more the longer he argues. So, the various anomalies in the dataset are somewhat interesting, but having weird outliers in the margins is an entirely expected effect. Just because there are not many datapoints. So when you filter for something marginal like Trump winning New Jersey, then the statistical error increases and therefore it is entirely unsurprising th…

> Just because there are not many datapoints. There are more than enough data points to determine the between-state error correlations, many of which seem to be very off. > Additionally, getting worked up about a 3% chance The weird between-state correlations actually have a large effect, they increase state and nationwide uncertainty and as a result Trump has a higher chance of winning.

It seems more likely that it's actually the other way. Nate Silver has specifically said he built the model to have relatively high uncertainty, especially with the volatility of this year, so this seems more like the outcome of intentional decisions to not let the model be overly confident.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#107
post #80

Earlier quoted context omitted.

Nate Silver said Trump had a 1 in 3 chance, which basically means one shouldn’t be surprised no matter the result. I’m not sure where this “all the polls were so far off in 2016!!” narrative comes from, but it’s wrong.

Nate was the outlier in that respect. But it’s true that the polls aren’t weren’t all that inaccurate in 2016: a bunch of important swing states were within the margin of error and Trump won some important states by very small margins. The mistake in 2016, IMO was a) the extrapolation that came from those polls and b) people paying way too much attention to national polls, which have very little connection to elector…

Silver's 2012 book "The Signal and the Noise" discusses our inability to rationally process probability, pointing out that commercial weather forecasts (e.g. Accuweather) never list a probably of rain under 20-25%. A 5% chance of rain is a mathematical possibility but people "feel" like 5% = "will never happen" and get angry if it rains.

Nobel Prize winner Daniel Kahneman's life's work is about this, what he calls "System 1" and "System 2" of our brain, where System 1 is a fast responder that provides insta-feedback but is largely incapable of processing mathematical inputs. His 2011 book "Thinking Fast and Slow" summarizes his work well.

I'm not sure popular media can be trained to frame statistical probabilities in a way that doesn't provide people with the certainty they crave. But who knows?

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#108

Earlier quoted context omitted.

Nowhere did I say it is. I simply think that, if one were following the right information, his win was not as unexpected as the coastal media presented it as being. Personally I would have put it about 60-40 Hillary-Trump.

You started off by calling the 2016 prediction a debacle , but now you're saying you would have put the odds about 12 percentage points differently. That doesn't seem like a big enough disagreement to warrant the kind of vehement criticism you're throwing around.

I called it a debacle because 99.9% of media sources, pundits, politicians, political figures, or anyone else thought that Hillary had anything less than a guaranteed win. I'm simply suggesting it was actually always a close race, but that the media ignored this because it went against their ideological model / they weren't familiar with places like the Rust Belt.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#109

Every single one of these models will break down this year. We are living in an unprecedented time. I can't understand how we can model how many people will vote, when we don't even know how many people have moved out of cities this year. Half of my friends have left San Francisco - if as many people left Philadelphia, The Twin cities, Milwaukee or Pittsburgh, then that really effects the outcome.

That's a neat point. My gut reaction is that there probably hasn't been enough people migrating to make a big difference. As in, I doubt enough people left CA to make it cut Republican and I doubt enough people moved to SC to make it cut Democrat. But that's just a gut reaction - I have no clue really.

It's a neat possibility to think about though. If there were enough people who did that, it would really depend on the demographics moving and where they're going. It could swing the election either way. I wonder if anyone has found numbers on this and attempted to model it.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#110
post #29
post #23

FiveThirtyEight has a page where you can choose winning states (condition on a certain outcome) and it will regenerate the prediction map, https://projects.fivethirtyeight.com/trump-biden-election-ma... This appears to be the what Andrew Gelman is also trying to do with their raw data. At the bottom of the 538 page it says, " If you choose enough unlikely outcomes, we’ll eventually wind up with so few simulations rem…

From the plots Andrew posted, it looks like the problem is not just sample size and that (some) individual state pairs have inverse correlations, e.g. https://statmodeling.stat.columbia.edu/wp-content/uploads/20...

I'd argue negative correlation on conditionals distributions can be reasonable here.

In that particular WA-MS example, if Trump suddenly took more liberal positions and somehow won WA (e.g., announces he's pro abortion), he would in fact be more at risk of losing Mississippi. The idea that these two states are in play already is fringe and would require some major idealogical (or other third variable) shifts.

Post reply on HN