Live data from Hacker News

FiveThirtyEight articles on the Internet Archive

fivethirtyeightindex.com

21–30 of 102 posts

Re: FiveThirtyEight articles on the Internet Archive

#21
Unfortunately most of the most important visualizations are broken in the archived version. Including the gun deaths visualization and I think the P-hacking interactive

https://web.archive.org/web/20230205124354/https://fivethirt...

It's kinda sad to know no one else will get to experience those interactive visualizations. Though its nice to see the approval comparison page still works

https://web.archive.org/web/20241031232233/https://projects....

Re: FiveThirtyEight articles on the Internet Archive

#22
post #6

For any, like myself, wondering " Who is Ben Welsh " ? Hello. My name is Ben Welsh. I'm an Iowan living in New York City. I am a reporter, an editor and a computer programmer. My job is to use those skills, together, to find and tell stories. I work at Reuters, the world's largest multimedia news provider, where I founded the organization's News Applications Desk. In that role, I lead the development of dashboards, d…

(Submitted title was "Ben Welsh made an index of all FiveThirtyEight articles on the Internet Archive" - we've since changed it)

Re: FiveThirtyEight articles on the Internet Archive

#23
post #22
post #6

For any, like myself, wondering " Who is Ben Welsh " ? Hello. My name is Ben Welsh. I'm an Iowan living in New York City. I am a reporter, an editor and a computer programmer. My job is to use those skills, together, to find and tell stories. I work at Reuters, the world's largest multimedia news provider, where I founded the organization's News Applications Desk. In that role, I lead the development of dashboards, d…

(Submitted title was "Ben Welsh made an index of all FiveThirtyEight articles on the Internet Archive" - we've since changed it)

Cheers for the clarity, that'll help me look less weird wrt above comment to future historians of archived HN threads :-)

TBH I enjoyed looking up Ben and finding out what he's about and done in the past far more than I did just knowing there's a 538 archive on IA.

Re: FiveThirtyEight articles on the Internet Archive

#24
post #23
post #22

Earlier quoted context omitted.

(Submitted title was "Ben Welsh made an index of all FiveThirtyEight articles on the Internet Archive" - we've since changed it)

Cheers for the clarity, that'll help me look less weird wrt above comment to future historians of archived HN threads :-) TBH I enjoyed looking up Ben and finding out what he's about and done in the past far more than I did just knowing there's a 538 archive on IA.

What do you think was perceived wrong with the old title?

Re: FiveThirtyEight articles on the Internet Archive

#26
post #24
post #23

Earlier quoted context omitted.

Cheers for the clarity, that'll help me look less weird wrt above comment to future historians of archived HN threads :-) TBH I enjoyed looking up Ben and finding out what he's about and done in the past far more than I did just knowing there's a 538 archive on IA.

What do you think was perceived wrong with the old title?

At a guess (my Telepathy/IP is weak today, I'm not reading dang at usual strength) .. the initially submitted title was "invented" for submission and didn't match the content title.

HN veers toward "the guts of the content w/out decoration" - limited additional information, framing, weasel words, perceived slanting, etc.

It's uncommon to name an author unless the author themself is an important part of "the story".

I personally have no issue with the original title, however it's not really for me (non US citizen) to judge whether the reporter in question has a name / identity that carries weight in US IT circles.

Re: FiveThirtyEight articles on the Internet Archive

#27
post #26
post #24

Earlier quoted context omitted.

What do you think was perceived wrong with the old title?

At a guess (my Telepathy/IP is weak today, I'm not reading dang at usual strength) .. the initially submitted title was "invented" for submission and didn't match the content title. HN veers toward "the guts of the content w/out decoration" - limited additional information, framing, weasel words, perceived slanting, etc. It's uncommon to name an author unless the author themself is an important part of "the story". I…

Insightful in spite of that difficulty :)

Re: FiveThirtyEight articles on the Internet Archive

#28
post #11

Earlier quoted context omitted.

Don't we need more than an index of Archive.org because whomever controls the domain could robots.txt these out of existence if they wanted to?

Archive.org mostly ignores robots.txt https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea...

People always use that link as reference to say that Internet Archive ignores robots.txt but it only actually says they are ignoring it for government sites. It suggests that they might do it for other sites in the future (of 2017), but does not actually say that that they have done it.

https://blog.archive.org/2018/04/24/addressing-recent-claims... which is a year later mentions that they have an automated process which is still following robots.txt for displaying old pages where the robots.txt was added later.

https://help.archive.org/help/using-the-wayback-machine/ does say they follow it for scraping, but this is phrased in such a way that would still be true for past sites whether or not they changed the policy. There is a page https://www.sysjolt.com/2021/archive-org-no-longer-honors-ro... which claims they don't follow it, but the site owner misspelled "robots" as "robot".

Re: FiveThirtyEight articles on the Internet Archive

#29

Love Ben but title can simply be: Index of FiveThirtyEight articles preserved by the Internet Archive

But that would be a false attribution. The Internet Archive did not create the index, Ben did. And the Internet Archive is not hosting the index, Ben is.

Ah, yes, could be worded better, fairplay. Point is the Ben attribution isn't needed in that place to avoid unnecessary confusion about who that is etc.

Re: FiveThirtyEight articles on the Internet Archive

#30
post #11

Earlier quoted context omitted.

Don't we need more than an index of Archive.org because whomever controls the domain could robots.txt these out of existence if they wanted to?

Archive.org mostly ignores robots.txt https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea...

The robots.txt file should be used to restrict (and, in some cases, slow down) crawling at the time it is being crawled, not for SEO or for restricting access to mirrors or for any other purpose. It should never apply retroactively. (Unfortunately it is sometimes used badly despite this.)
Post reply on HN