Live data from Hacker News

FiveThirtyEight has a GitHub repo with story-related data and scripts

github.com

21–30 of 38 posts

Re: FiveThirtyEight has a GitHub repo with story-related data and scripts

#21
post #11

No story about Nate Silver and 538 is complete without mentioning his new rival in the prediction space - Carl Diggler - who has been kicking his butt on the presidential primaries, calling tough elections correctly where Nate has refused to give more than a "Maybe Sanders, Maybe Clinton..." prediction. Carl is a work of fiction, invented by a couple of journalists/friends as a parody of pundits: https://www.washingt…

538 is the classic "fighting the last war" problem. In 2012 the issue was that polling was unreliable, it was difficult to figure out what was actually going on with the electorate. FiveThirtyEight figured out how to crack that problem, by applying number crunching. This cycle, the problem is that the electorate hasn't actually decided what it wants yet, and there statistical methods aren't helpful at all. There bein…

Was 538 that good at predicting outcomes of the 2012 primaries? I know it's well established that they did a good job on the 2008 and 2012 general elections.

Primaries have to be harder because there is far less polling information compared to the general election. Yes, they were unable to predict the Trump victory, but if the get 90-100% of states correct in the general election, then I doubt that much has changed with how polling reflects the electorate.

Re: FiveThirtyEight has a GitHub repo with story-related data and scripts

#22
post #11

No story about Nate Silver and 538 is complete without mentioning his new rival in the prediction space - Carl Diggler - who has been kicking his butt on the presidential primaries, calling tough elections correctly where Nate has refused to give more than a "Maybe Sanders, Maybe Clinton..." prediction. Carl is a work of fiction, invented by a couple of journalists/friends as a parody of pundits: https://www.washingt…

> We called 19 out of the past 19 contests. FiveThirtyEight, whose model cannot work without polling, accurately predicted 13.

FiveThirtyEight's model could work without polling (it's not the only predictor they use), but they don't publish results until there are at least two polls conducted. I would call that a feature, not a bug.

Also, FiveThirtyEight correctly predicted the outcome of every primary race that it modeled, on both sides, except one (Michigan Democratic primary).

Re: FiveThirtyEight has a GitHub repo with story-related data and scripts

#23
post #18

I picked a project at random -- the Tarantino one, which lists every death or curse word in every Tarantino movie -- and browsed through the data: https://github.com/fivethirtyeight/data/blob/master/tarantin... I found it interesting that the author was willing to catalog each curse word by writing it in its entirety -- except one, the n-word (which, by the way, appears 179 times in Tarantino's movies, with the bulk…

n-word -> nigger and f-word -> fuck I understand, but what is the w-word? I really don't get this obsession over hiding the existence of these words. They obviously don't exist in isolation and are a symptom of a different problem. Perhaps by hiding them people seek to pretend the world has solved these problems? On top of that, their usage as negative words is only propagated by these shortenings. Words like nigger…

W-word is "wetback", I think. I've never heard anyone use it.

Re: FiveThirtyEight has a GitHub repo with story-related data and scripts

#24
post #18

I picked a project at random -- the Tarantino one, which lists every death or curse word in every Tarantino movie -- and browsed through the data: https://github.com/fivethirtyeight/data/blob/master/tarantin... I found it interesting that the author was willing to catalog each curse word by writing it in its entirety -- except one, the n-word (which, by the way, appears 179 times in Tarantino's movies, with the bulk…

n-word -> nigger and f-word -> fuck I understand, but what is the w-word? I really don't get this obsession over hiding the existence of these words. They obviously don't exist in isolation and are a symptom of a different problem. Perhaps by hiding them people seek to pretend the world has solved these problems? On top of that, their usage as negative words is only propagated by these shortenings. Words like nigger…

Um, that link establishes that the word has had very strongly negative connotations since at least the 1850s.

The article states that it had "more neutral" connotations earlier on...in a society that enforced racial human enslavement.

Re: FiveThirtyEight has a GitHub repo with story-related data and scripts

#25
post #18

I picked a project at random -- the Tarantino one, which lists every death or curse word in every Tarantino movie -- and browsed through the data: https://github.com/fivethirtyeight/data/blob/master/tarantin... I found it interesting that the author was willing to catalog each curse word by writing it in its entirety -- except one, the n-word (which, by the way, appears 179 times in Tarantino's movies, with the bulk…

n-word -> nigger and f-word -> fuck I understand, but what is the w-word? I really don't get this obsession over hiding the existence of these words. They obviously don't exist in isolation and are a symptom of a different problem. Perhaps by hiding them people seek to pretend the world has solved these problems? On top of that, their usage as negative words is only propagated by these shortenings. Words like nigger…

> what is the w-word?

> I really don't get this obsession over hiding the existence of these words.

You've kind of answered your own question there. If you hide the words, then people who don't know them (eg children), don't learn a new offensive term.

I'm not saying I support the policy; I'm just saying that I understand why people do it.

Re: FiveThirtyEight has a GitHub repo with story-related data and scripts

#26
post #11

No story about Nate Silver and 538 is complete without mentioning his new rival in the prediction space - Carl Diggler - who has been kicking his butt on the presidential primaries, calling tough elections correctly where Nate has refused to give more than a "Maybe Sanders, Maybe Clinton..." prediction. Carl is a work of fiction, invented by a couple of journalists/friends as a parody of pundits: https://www.washingt…

From the article you linked:

> Even when all of Silver’s models for a given race turn up wrong, it never seems to be FiveThirtyEight’s fault.

That is not at all a fair summary of 538's position. They always take responsibility when they come up short.

> When the site badly whiffed on last year’s British election, it was the pollsters who erred. (link to: http://fivethirtyeight.com/datalab/what-we-got-wrong-in-our-...)

The article they link to here doesn't shift blame to the pollsters at all. While they note that the polls were wrong, they take full responsibility for their bad predictions:

"It’s our job as forecasters to report predictions with accurate characterizations of uncertainty, and we failed to achieve that in this election. We are not trying to make excuses here; we are trying to understand what went wrong."

It's a little bit annoying seeing some people who are satirizing the media class unapologetically misrepresenting others in the same breath.

Re: FiveThirtyEight has a GitHub repo with story-related data and scripts

#27
post #11

No story about Nate Silver and 538 is complete without mentioning his new rival in the prediction space - Carl Diggler - who has been kicking his butt on the presidential primaries, calling tough elections correctly where Nate has refused to give more than a "Maybe Sanders, Maybe Clinton..." prediction. Carl is a work of fiction, invented by a couple of journalists/friends as a parody of pundits: https://www.washingt…

As fun as the story of Carl Diggler is, his work is statistically insignificant. The podcast Reply All (which you should be listening to) asked a professor who studies the science of predictions whether the Diggler guys are onto something [1].

They are not. The professor wanted some kind of statistically significant proof that they were any better than Paul, the octopus who successfully predicted an astonishingly high number of soccer matches by random chance [2].

Say what you will about 538, but at least they care about not using anecdotal or imagined evidence. At least they're rigorous with their methodology. They're also actually good at it, unlike most pundits [3].

[1]: https://gimletmedia.com/episode/58-earth-pony/

[2]: https://en.wikipedia.org/wiki/Paul_the_Octopus

[3]: https://en.wikipedia.org/wiki/FiveThirtyEight#Final_projecti...

Re: FiveThirtyEight has a GitHub repo with story-related data and scripts

#28

Earlier quoted context omitted.

n-word -> nigger and f-word -> fuck I understand, but what is the w-word? I really don't get this obsession over hiding the existence of these words. They obviously don't exist in isolation and are a symptom of a different problem. Perhaps by hiding them people seek to pretend the world has solved these problems? On top of that, their usage as negative words is only propagated by these shortenings. Words like nigger…

> what is the w-word? > I really don't get this obsession over hiding the existence of these words. You've kind of answered your own question there. If you hide the words, then people who don't know them (eg children), don't learn a new offensive term. I'm not saying I support the policy; I'm just saying that I understand why people do it.

> children don't learn a new offensive term.

Um, yes, yes they do. Obscuring words is to protect you from the reminder of the terrible history human kind has had toward people that look different from you. It doesn't protect a child's innocence.

Re: FiveThirtyEight has a GitHub repo with story-related data and scripts

#29

Data is nice, but even with it, 538 has gone alarmingly down the path of clickbait lately. Especially Casselman, with smarmy headlines like "The Rising Unemployment Rate Is Good News" or "Stuck In Your Parents’ Basement? Don’t Blame The Economy".

It's not lately. There was a precipitous drop in quality (of math and subject) as soon as they went to ESPN.

Re: FiveThirtyEight has a GitHub repo with story-related data and scripts

#30
post #27
post #11

No story about Nate Silver and 538 is complete without mentioning his new rival in the prediction space - Carl Diggler - who has been kicking his butt on the presidential primaries, calling tough elections correctly where Nate has refused to give more than a "Maybe Sanders, Maybe Clinton..." prediction. Carl is a work of fiction, invented by a couple of journalists/friends as a parody of pundits: https://www.washingt…

As fun as the story of Carl Diggler is, his work is statistically insignificant. The podcast Reply All (which you should be listening to) asked a professor who studies the science of predictions whether the Diggler guys are onto something [1]. They are not. The professor wanted some kind of statistically significant proof that they were any better than Paul, the octopus who successfully predicted an astonishingly hig…

The point is that even with 'rigorous methodology' and 'evidence' 538 doesn't do as well as Diggler.

Predicting elections is, for the most part, an easy task. Polls are usually good enough to get some picture of the election at hand. One major shortcoming of a primarily evidence-based approach is how to call races without much polling (small sample size). Silver has consistently refused to cover these races at all.

But Silver commits two sins that you ignore: he wants to show off his 'accuracy' and he wants to improve on the polls. These two problems kind of intertwine - the 538 'polls-plus,' which incorporates predictors Silver thinks are significant e.g. endorsements or Nate's gut, performs worse than the polls forecast, which is just a weighted average of recent state polls. What kind of rigorous methodology allows for two different predictions and then picking the best one, after the fact? What kind of rigorous methodology results in a prediction method that's shown to be worse than reading the newspaper, time and time again? Where's the evidence behind endorsements working this year?

As a side note, Philip Tetlock is the Wharton professor, and Dan Gardner (interviewed here) is some dude.

Post reply on HN