Live data from Hacker News

Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

statmodeling.stat.columbia.edu

161–170 of 243 posts

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#161

Andrew Gelman designed the 538 model in 2007. Nate Silver authored an adjustment to polls used in that model. Polls have more impact if they are more representative of statewide turnout among demographic things he chose like “black” and “low income.” This is why his predictions were so accurate for Obama’s 2008 and 2012 elections, and likely why they were so inaccurate in 2016. Gelman’s own grad student is the only p…

Nate Silver said Trump had a 1 in 3 chance, which basically means one shouldn’t be surprised no matter the result. I’m not sure where this “all the polls were so far off in 2016!!” narrative comes from, but it’s wrong.

It was more like 1 in 5.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#162

Earlier quoted context omitted.

Nit: Have they counted for the possibility of a tie? US elections allow for a tie in the Electoral College (which then kicks off a supremely strange and legalistic process).

Before commenting you could at least search for the word "tie": The probability of an electoral-college tie is

Not everyone can read that little grey font as well as the next person. I guess I apologize for my poor eyesight then.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#163
post #115
post #28

Not sure if it’s wrong to put this here, but here is a link to their election forecast. https://projects.economist.com/us-2020-forecast/president You can compare this to the 538 model and see where these two teams and forecasts disagree.

Can someone help me understand what odds like this mean in the context of an election? The model says that Trump has a 1 in 10 chance of winning. With a fair 10-sided die it makes sense that you have a 1 in 10 chance of any given side rolling face up. But what is the die that is being rolled in these election statistics? What is the "chance" element that is being predicted?

That's a good question and it's not clear. As someone mentioned, here "chance" includes both uncertainty (facts that we don't know), and randomness of nature (things that will happen in the future that cannot be deterministically deduced from the state of the world today). Depending on your philosophy these may overlap. Next, someone mentioned Bayesian vs frequentist.

The frequentist interpretation is roughly that if I go around making my best possible predictions, and we lump together all the things that I predict at 10%, about 1 in 10 of those things happen and the rest don't. But I wouldn't be able to be more specific about which ones in that group are more likely than others.

The Bayesian interpretation is that I can really view the world as flipping coins -- I don't care whether it's due to my lack of knowledge or "true" randomness -- and as far as I can tell, the coin flip involved here is 1 in 10.

We can also use a gambling interpretation. Here's one based on security of python's random module. Imagine the following three lotteries I offer you. In lottery A, you get $100 if Trump is elected. In lottery B, you get $100 if the following python code returns true on my laptop:

    random.random() 
In lottery C, you get $100 if this code returns true:

    random.random() 
If you would rather have lottery A than B, and you'd rather have C than A, then in some sense that you believe Trump has a 1 in 10 chance.

Now there's an interesting extra layer to all of this because it's a model predicting, not a person. In a short space, I would basically say that we've trained models to predict in ways that are not inconsistent with any of the interpretations above, when put into situations where that is testable. Then we use them in situations where it might not be, like this.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#164
post #115
post #28

Not sure if it’s wrong to put this here, but here is a link to their election forecast. https://projects.economist.com/us-2020-forecast/president You can compare this to the 538 model and see where these two teams and forecasts disagree.

Can someone help me understand what odds like this mean in the context of an election? The model says that Trump has a 1 in 10 chance of winning. With a fair 10-sided die it makes sense that you have a 1 in 10 chance of any given side rolling face up. But what is the die that is being rolled in these election statistics? What is the "chance" element that is being predicted?

It pretty much means nothing. These sort of models produced wrong result again and again. For me the biggest question mark if that if we know that recent (last 5) elections were very close how can you predict somebody winning with 93% chance? Maybe I do not understand something here.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#165
post #161

Earlier quoted context omitted.

Nate Silver said Trump had a 1 in 3 chance, which basically means one shouldn’t be surprised no matter the result. I’m not sure where this “all the polls were so far off in 2016!!” narrative comes from, but it’s wrong.

It was more like 1 in 5.

It's still up:

https://projects.fivethirtyeight.com/2016-election-forecast/

28.6% is between 1/3 and 1/4, definitely not 1/5.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#166

Earlier quoted context omitted.

> Just because there are not many datapoints. There are more than enough data points to determine the between-state error correlations, many of which seem to be very off. > Additionally, getting worked up about a 3% chance The weird between-state correlations actually have a large effect, they increase state and nationwide uncertainty and as a result Trump has a higher chance of winning.

It seems more likely that it's actually the other way. Nate Silver has specifically said he built the model to have relatively high uncertainty, especially with the volatility of this year, so this seems more like the outcome of intentional decisions to not let the model be overly confident.

I think it's an error in the model structure. If their goal was to artificially increase uncertainty, there are more reasonable ways to do that than adding weird between-state correlations (like the uncertainty index which is part of the model). WA and MS definitely should not have a large negative correlation.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#167
post #35

Earlier quoted context omitted.

They said Trump had a 1 in 4 chance. That's very high. NYT had something like 1 in 20 chance for Trump.

That doesn't really answer my question. It only indicates that they were slightly less wrong than every other media source, not that they have a good model. If I had a laptop that only worked 1/4th of the time, rather than 1/20th of the time, would that make it a reliable laptop? I don't think so.

If everyone was wrong, it is reasonable to believe a low probability event occurred. However, given the extent to which people predicted a Trump loss (say 1 in 20), which is significantly rarer, given that the event occurred it suggests the model that predicted a Trump win with the greatest probability to likely be a more accurate model.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#168

Earlier quoted context omitted.

> If you fit a model and don't explicitly constrain against "un-physical" results like negative correlations, you'll end up with them. The Economist model does exactly that, and all of their correlations are positive. I recommend reading their methodology, they know what they're doing (I wouldn't say the same about 538). Andrew Gelman has developed some of the Bayesian methods and software that people like Nate Silve…

I think the question is if it matters to the predictive accuracy of the model. Just because it puts out results you can't envision actually happening on the margins doesn't mean they can't happen, or that they can't be valuable in presenting a holistic result. It's clear that the models are tuned differently, but from Silver's replies in the PS's, it seems that he's ok with these artifacts being part of the model.

Yes, it increases the state-level and national uncertainty intervals (Andrew Gelman has talked about this several times on his blog), which improves Trump's odds.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#169
post #89

To me, this mostly tracks with what 538 has said on record about how their model works and the design philosophy behind parts of it. To me, what Nate means when he says "directionally the right approach in terms of our model's takeaways" is that these sorts of wild and unintuitive outcomes are part of the point of the way the model is constructed. Specifically, that when you get off into the weird situations like Tru…

Yeah I think for Trump to win washington state, he'd have to do something to appeal to voters there in a way that would likely cause his red state base to abandon him. The negative correlation makes sense when we think about how difficult it is for everyone in Washington to suddenly turn conservative and everyone in Mississippi to turn liberal. Much more likely is that the crazy thing is that the candidate or circums…

I think there's nothing Trump can do to win Washington; rather, more accurate to say there are many things Biden could do to lose Washington.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#170

Earlier quoted context omitted.

Saying that the ultimate outcome had a 1 in 4 chance is not wrong, slightly wrong, or less wrong. If the weatherman says there's a 1 in 4 chance of rain, and it rains, he wasn't wrong.

To take it a step further, if that was the forecast 4 days in a row, you would expect it to rain one of those days.

Not necessarily.

There is a 0.750.750.75*0.75 or a 31% chance of no rain at all

Post reply on HN