Live data from Hacker News

Link rot and content drift are endemic to the web

theatlantic.com

21–30 of 220 posts

Re: Link rot and content drift are endemic to the web

#21

By the way, the technical side of this is very interesting. If you look at the tools mentioned (the wayback machine, but also perma.cc and other archival solutions), almost all of them rely on a single semi-modern tech stack that produces WARCs (web archives - ISO - ISO 28500:2017 https://iipc.github.io/warc-specifications/specifications/wa... ). The main crawler still seems to be heritrix3 ( https://github.com/inter…

Honestly, it would be a better use of surplus resources than crypto mining.

If only there are a way to algorithmically tie a proof of work for a new cryptocurrency to archival of the internet in a way that wouldn't be easily gamed (by people archiving easy to access content or highly redundant archival of trivia).

Re: Link rot and content drift are endemic to the web

#22

yes. i am down to HN and Reddit. Google calendar and telegram. I don't even know how to find cool stuff online anymore. Google SERPs are all business driven now unless you're research news.

The article is about link rot, not cultural rot.

And it seems to overlook a pretty straightforward question: in the era of the search engine, how much of an issue is link rot?

I've hit bad links before. Four out of five times, I can do a general search for the title of the document that should have been at the link or the quoted excerpt that the document I'm reading pulled from the link, and I get a clone of the document posted somewhere else.

Re: Link rot and content drift are endemic to the web

#23
post #5

yes. i am down to HN and Reddit. Google calendar and telegram. I don't even know how to find cool stuff online anymore. Google SERPs are all business driven now unless you're research news.

All the cool stuff has gone back to IRL. Once people stopped making potato gun websites, the internet really stopped blossoming into an amazingly vibrant space. I would recommend hackaday.com because it hasn't changed in quite some time.

FWIW, Siemens recently bought hackaday (or well, Siemens bought a company called Supplyframe which was the owner of Hackaday), so lets see how long that will last..

Re: Link rot and content drift are endemic to the web

#25

By the way, the technical side of this is very interesting. If you look at the tools mentioned (the wayback machine, but also perma.cc and other archival solutions), almost all of them rely on a single semi-modern tech stack that produces WARCs (web archives - ISO - ISO 28500:2017 https://iipc.github.io/warc-specifications/specifications/wa... ). The main crawler still seems to be heritrix3 ( https://github.com/inter…

Honestly, it would be a better use of surplus resources than crypto mining. If only there are a way to algorithmically tie a proof of work for a new cryptocurrency to archival of the internet in a way that wouldn't be easily gamed (by people archiving easy to access content or highly redundant archival of trivia).

https://spec.filecoin.io/algorithms/pos/

Re: Link rot and content drift are endemic to the web

#26

yes. i am down to HN and Reddit. Google calendar and telegram. I don't even know how to find cool stuff online anymore. Google SERPs are all business driven now unless you're research news.

I wish there was something like Reddit that could be organized by topic, but had the simplicity of HN's design instead of the monstrocity Reddit has become. My guess is it would still succumb to the Reddit Hive Mind effect without a reasonably benevolent moderation team though. For all the times I've said that HN basically does the same thing, I have to admit that it is much better about keeping it in check.

Reddit has been pushing to be like other social media sites now. They've added profile pictures and avatars, not to mention they're pushing video content like nothing else. They've got livestreaming and try to saturate your front page with as much video as possible. They've also changed the way their app handles video links to be more like TikTok or YouTube.

It used to be a lot like HN - discussions around links to articles. I wish there was a community with the feel of HN with the wide net of Reddit.

It feels like all social media is converging; Snapchat, Instagram, Facebook, Reddit, TikTok, Youtube, all an endless stream of ai-curated short videos that you can swipe through over and over.

Re: Link rot and content drift are endemic to the web

#27

yes. i am down to HN and Reddit. Google calendar and telegram. I don't even know how to find cool stuff online anymore. Google SERPs are all business driven now unless you're research news.

You either have to visit aggregating sites that have likeminded people (HN/Reddit/Obscure FB groups/Group chats/Forums) or know exactly what you're looking for.

Re: Link rot and content drift are endemic to the web

#28
post #5

yes. i am down to HN and Reddit. Google calendar and telegram. I don't even know how to find cool stuff online anymore. Google SERPs are all business driven now unless you're research news.

All the cool stuff has gone back to IRL. Once people stopped making potato gun websites, the internet really stopped blossoming into an amazingly vibrant space. I would recommend hackaday.com because it hasn't changed in quite some time.

> Once people stopped making potato gun websites, the internet really stopped blossoming into an amazingly vibrant space.

Have they or have you stopped looking? I'd say the former.

Re: Link rot and content drift are endemic to the web

#29

By the way, the technical side of this is very interesting. If you look at the tools mentioned (the wayback machine, but also perma.cc and other archival solutions), almost all of them rely on a single semi-modern tech stack that produces WARCs (web archives - ISO - ISO 28500:2017 https://iipc.github.io/warc-specifications/specifications/wa... ). The main crawler still seems to be heritrix3 ( https://github.com/inter…

WARC can record and replay single-page apps, but it struggles with knowing where a "page" begins and ends.

There was a time when I was furious with the web going to hell and I investigated the possibility of "web without browsers" that started with making a WARC capture of page and putting pages through extensive filtering and classification before the user sees anything.

With interactive capturing you can push a button to indicate that a page is done "loading" but with automated capturing you can't really know that the page is done or that you got a good capture. That ended the project right there.

Re: Link rot and content drift are endemic to the web

#30
It's an Atlantic article so it's long, and several of the comments here show that people aren't actually reading the whole thing... but I did and it's worth the time. It's not only about links being dead, it's about the lack of transparency & audit when content is changed via takedown requests, it's about dead links showing up in decades-old supreme court decisions, it's about private industry's lack of incentive for wanting to improve any of these issues, etc.
Post reply on HN