Live data from Hacker News

Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

statmodeling.stat.columbia.edu

81–90 of 243 posts

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#81
post #8

> It didn't take very long to do the analysis. But it did then take another hour or so to write it up. It's very interesting to see how long it takes people to do things. I am amazed that entire article took 1 hour to type up. I've spent entire afternoons trying to write shallower pieces of work.

[deleted]

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#82
post #35

Earlier quoted context omitted.

They said Trump had a 1 in 4 chance. That's very high. NYT had something like 1 in 20 chance for Trump.

That doesn't really answer my question. It only indicates that they were slightly less wrong than every other media source, not that they have a good model. If I had a laptop that only worked 1/4th of the time, rather than 1/20th of the time, would that make it a reliable laptop? I don't think so.

If you have a die that rolls a one 1/6 of the time, do you consider the die wrong?

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#83
post #69

Funny thing, can't submit this to r/politics, as they seem to have a tightly curated whitelist of allowed domains that must "Be notable, as defined by our domain notability guidelines. Notable domains will consist of news organizations, research organizations, political advocacy groups, governmental agencies / bodies, and political parties." And apparently columbia.edu does not fulfill those criteria.

Don't bother, /r/politics is one of the most censored places on the internet. It's best avoided. Despite its description don't expect any actual adult discussion of politics there.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#84
post #45

It seems like the behavior between WA and MS could just be statistics saying that WA and MS always[1] vote for the opposite candidate, rather than considering a massive sudden change in the direction that one of them votes in. E.g. it's not reflecting who they vote, just who they most vehemently disagree with. I'm not sure why that kind of interstate correlation should impact predictions? IANAS but it feels like thes…

> I'm not sure why that kind of interstate correlation should impact predictions?

538 has low positive correlations between states on average, which actually has a big impact, it increases overall uncertainty (and therefore Trump's win probability). Why? If the states are not correlated, you usually end up with a few states going off the rails, like Trump winning Colorado without any nationwide swing.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#85
post #69

Funny thing, can't submit this to r/politics, as they seem to have a tightly curated whitelist of allowed domains that must "Be notable, as defined by our domain notability guidelines. Notable domains will consist of news organizations, research organizations, political advocacy groups, governmental agencies / bodies, and political parties." And apparently columbia.edu does not fulfill those criteria.

Did you try asking the moderators to add it to the whitelist?

In a way. They have a form for suggesting new domains to the whitelist: https://goo.gl/forms/lRQikA1rI0bVbKCl1

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#86

Earlier quoted context omitted.

That doesn't really answer my question. It only indicates that they were slightly less wrong than every other media source, not that they have a good model. If I had a laptop that only worked 1/4th of the time, rather than 1/20th of the time, would that make it a reliable laptop? I don't think so.

If you have a die that rolls a one 1/6 of the time, do you consider the die wrong?

No, but I don't consider it useful toward predicting the outcome, which is of course what this is all about.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#87
post #15

Earlier quoted context omitted.

No one is purporting election forecasting to be scientific.

Some of your sibling comments seem to suggest otherwise.

Well, HN commenters don’t speak for the people doing election forecasting. I think Silver would describe it as educated guesswork.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#88

Every single one of these models will break down this year. We are living in an unprecedented time. I can't understand how we can model how many people will vote, when we don't even know how many people have moved out of cities this year. Half of my friends have left San Francisco - if as many people left Philadelphia, The Twin cities, Milwaukee or Pittsburgh, then that really effects the outcome.

That's an interesting thought, but San Francisco aside, it seems like most people moving out of cities are just moving to the suburbs of those cities, which shouldn't affect the presidential election calculus?

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#89
To me, this mostly tracks with what 538 has said on record about how their model works and the design philosophy behind parts of it. To me, what Nate means when he says "directionally the right approach in terms of our model's takeaways" is that these sorts of wild and unintuitive outcomes are part of the point of the way the model is constructed.

Specifically, that when you get off into the weird situations like Trump winning Washington state, it's likely something incredibly weird has happened - something that likely has no historical precedent, so it may actually be a more sane thing to do to assume that now almost everything is backwards and Biden would win a bunch of states he shouldn't either.

To me, this points to a general willingness in the 538 model to just go "who knows" and build in some room for insane things to happen on the fringes. The Friday podcast episode about the 538 model specifically mentions that they have large/fat tails on their distribution that make it nearly impossible for someone to get over 95% chances of winning on a national level, and these sorts of wild results seem like the outcome of that. If you bake in an assumption that there's always a 5% chance of something crazy happening, that chance has to come from something in the data somewhere that reflects the ability of that to happen numerically, and thus will have numerical outcomes that seem impossible.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#90
> I’d think that if Trump were to win New Jersey or, even more so, California, that this would most likely happen only as part of a national landslide of the sort envisioned by Scott Adams or whatever.

That's a valid intuition to have but you can also clearly make the argument that if Trump wins California you're in such a weird scenario that using the traditional wisdom about correlation is dangerous. The point that 538 have tried repeatedly to make is that firstly: if you're conservative in your level of confidence you'll give a higher likelihood to outliers, and secondly: It's not particularly useful to focus on whether X has a 3% or 4% chance.

If Trump wins California, we aren't going to be talking about whether the chance was 3% or 0.3% we're going to be talking about that Nuclear explosion that wiped out 25million Californians.

For the same logic the reason that Trump winning Alaska given winning New Jersey is lower than given losing New Jersy is because your sample size is rubbish. The chance of Trump winning Alaska given losing New Jersey is an accurate number, the number of Trump winning Alaska given winning New Jersey is like saying "How likely is it Trump wins Alaska given the UK gains US statehood" it's like.... well... if that happens then we're so far outside of what the model thinks can happen then you should be that we're just gonna say it's 50:50 - because who the hell knows.

It's not like saying "Oh well if X swing state goes blue, Y will probably follow", the scenarios in this article are so bizarre that the model should rightly be very cautious and probably default to either refusing to give an answer or just default to 50:50 or the same probably ignoring that data. The implicit bias in this analysis seems to be that if NJ went Red that would be because Trump won by a big margin, but that's not a likely enough scenario to actually get numbers for, and is so unlikely that things like "The supreme court threw out all the ballots for inner city areas" start to become valid possibilities.

Post reply on HN