Live data from Hacker News

Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

statmodeling.stat.columbia.edu

141–150 of 243 posts

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#141
post #12

Earlier quoted context omitted.

Yeah, but what if we could never roll that particular die again? I think that's what he's talking about.

Yes, which is why it’s foolish to talk about the outcome of that single race. You are right that there will never be another 2016 election between Clinton and Trump. However, that is only one of hundreds of forecasts made by 538 across multiple election cycles, so we can see how often their probabilistic outcomes align with actual outcomes. The fact that their track record is fairly good across all of these is eviden…

Like I told the other guy, I'm not defending his position, I'm stating it in different terms. And stating why pointing out the odds of a fair die isn't a good counter argument.

Grandposter doesn't believe the die is fair. That's a different argument than the guy I responded to made.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#142
post #121

Earlier quoted context omitted.

I'm not trying to get points, or keep points. I'm just trying to point out that the models are not going to work this year and its crazy to assume they will.

The models will work just fine. We are talking about large numbers here, and your small sample of anecdotes does not mean anything. No, there is no mass exodus. No, it is not going to change the results. Yes, the current polls are more accurate than just about any in history and we have an abundance of high-quality state level polls to back up the predictions. If you want to know where the models are going to break d…

"Garbage in, Garbage out".

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#143

Earlier quoted context omitted.

Now that you bring that up, if we had a 100% accurate poll that would be really good for productivity wouldn't it? Perhaps it wouldn't give voters the same feeling of self-determination but it'd save a lot of resources in fundraising, going out to vote, counting votes

I mean technically the election is a 100% accurate poll. You get data from every voter. To run a 100% accurate poll would require you to sample every voter, so it would literally be an election.

> I mean technically the election is a 100% accurate poll. You get data from every voter.

Actually an election is literally a poll in the sense of a sample, you are counting people at "polling stations" in order to gauge the public mood about who should be president.

When you see it that way, Nate Silver is predicting a sample of an unknown distribution, the "true" distribution of people's preferences.

The fact that you can only make one officially binding sample ("the voters") is a practicality, as is the fact that there's an electoral college that means votes have different values. The fact that turnout matters is another issue, a statistician might call it a sampling problem.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#144
post #115
post #28

Not sure if it’s wrong to put this here, but here is a link to their election forecast. https://projects.economist.com/us-2020-forecast/president You can compare this to the 538 model and see where these two teams and forecasts disagree.

Can someone help me understand what odds like this mean in the context of an election? The model says that Trump has a 1 in 10 chance of winning. With a fair 10-sided die it makes sense that you have a 1 in 10 chance of any given side rolling face up. But what is the die that is being rolled in these election statistics? What is the "chance" element that is being predicted?

The odds for a face on a d10 would be 1 in 9 (1:9).

It's different than probability (1/10 = .1)

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#145
I disagree with the author on the idea that tail is too fat for isolated anomalies. There are most certainly events that can happen, which may lead to a red California or a blue Alabama.

Presidential assassination, war, video proof of something incredibly heinous (pedophilia?), etc. can absolutely lead to these outcomes. You don't even have to go that far back. Nixon and Reagan flipped states like no-one's business.

I do however agree, that 538's state-state correlation model seems weak.

California and Alabama would only flip during a wave, and that wave would consume any and all states. The fact that 538's model doesn't strongly show that pattern is a failing of it. But, it is not clear if a model that inaccurately models the unlikeliest of events (california flipping while Florida stays blue), does not necessarily mean that it is terrible predictor of it's primary target (Presidential likelihoods).

As a data scientist, I can totally understand Nate's hesitation. Do you impose strong priors on the model to reflect strong domain intuition or do build a model that best characterizes the data it is based on. In the presence of infinite data, you should abandon all domain based priors. For single digit data points, priors are essential. For any number of data in between, it is anyone's best guess.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#146

Andrew Gelman designed the 538 model in 2007. Nate Silver authored an adjustment to polls used in that model. Polls have more impact if they are more representative of statewide turnout among demographic things he chose like “black” and “low income.” This is why his predictions were so accurate for Obama’s 2008 and 2012 elections, and likely why they were so inaccurate in 2016. Gelman’s own grad student is the only p…

Nate Silver said Trump had a 1 in 3 chance, which basically means one shouldn’t be surprised no matter the result. I’m not sure where this “all the polls were so far off in 2016!!” narrative comes from, but it’s wrong.

It comes from innumerate journalism, and an innumerate population. Next time someone laughs off being bad at math, you should point out that being unable to read is no laughing matter, and being unable to understand numbers shouldn't be either.

The only sensible way to predict probabilities that aren't extreme is to tell people how the model works and the figures it is currently spitting out. That's is the great thing about these kinds of blog posts, people are kicking the tyres, not just looking at the car.

Nobody predicting a one-off election with a rather special candidate would summarize a 33% chance as equivalent to having no chance.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#147
For a post on a blog presumably focused on statistics the author seems oddly concerned with a model that predicts odd events as being "possible". The difference between "possible" and "impossible" is \eps Likewise asserting that California with a 3% chance of going Trump is absurd is an unreasonable degree of overconfidence. Assuming maximizing expected return, the author is implying that they would be willing to take a bet that Trump would lose California with odds >> 97::3, i.e. presumably they would take a bet where I bet $1 to every $99 they bet. To be critical of a model based on outcomes it predicts with tiny probability you need truly remarkably biased priors.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#148
post #115
post #28

Not sure if it’s wrong to put this here, but here is a link to their election forecast. https://projects.economist.com/us-2020-forecast/president You can compare this to the 538 model and see where these two teams and forecasts disagree.

Can someone help me understand what odds like this mean in the context of an election? The model says that Trump has a 1 in 10 chance of winning. With a fair 10-sided die it makes sense that you have a 1 in 10 chance of any given side rolling face up. But what is the die that is being rolled in these election statistics? What is the "chance" element that is being predicted?

You're comparing two scenarios, one in which you know all the facts, and one in which you don't.

In the dice toss scenario, we know everything relevant. In the election scenario, we don't.

A model like this is attempting to say "these are the rules we think exist. Based on the rules, and assuming the data is off by some random distribution, here's what we think could happen".

What different forecasters disagree about is what the rules are. For example, the relevance of certain demographic characteristics and the potential variance between polling (conducted prior to the election) and actual election results.

There's a huge amount of assumptions, and forecasters disagree on those assumptions. We have very little historical data (polling is very recent) and even with complete historical data, future elections do not always conform to past elections.

Re: Reverse-engineering the problematic tail behavior of Fivethirtyeight forecast

#150
post #115
post #28

Not sure if it’s wrong to put this here, but here is a link to their election forecast. https://projects.economist.com/us-2020-forecast/president You can compare this to the 538 model and see where these two teams and forecasts disagree.

Can someone help me understand what odds like this mean in the context of an election? The model says that Trump has a 1 in 10 chance of winning. With a fair 10-sided die it makes sense that you have a 1 in 10 chance of any given side rolling face up. But what is the die that is being rolled in these election statistics? What is the "chance" element that is being predicted?

It's less about pure random chance, and more about our uncertainty. Compare it to a weather forecast that says there's a 10% chance of rain tomorrow. In the same way that weather forecasts get better over time (better atmospheric measurements, more sophisticated computer models), we could potentially do more to measure what the outcome of the election will be. And it might be theoretically possible (albeit highly unrealistic) to predict it with complete accuracy, given enough data. But we're not in that situation, hence uncertainty.

(There are a couple of caveats about election forecasting as opposed to weather forecasting. The first is the "October surprise," a sudden revelation that changes the election. This cycle, it was arguably Trump's covid diagnosis, although that tended if anything to push the results further in the direction they seemed to be going on their own, rather than upset any trend. The second is that, unlike with weather systems, measuring voter behavior (and widespread reporting on these measures) can change people's behavior. The effect of this is hotly contested, but one of the many explanations of Trump's victory in 2016 which hinged on turnout in a few key states is that those states were predicted wins for Clinton, so Clinton voters didn't bother voting. Despite occasional jokes to the contrary, it doesn't rain just to spite the weatherman.)

Post reply on HN