Live data from Hacker News

Link rot and content drift are endemic to the web

theatlantic.com

11–20 of 220 posts

Re: Link rot and content drift are endemic to the web

#11
post #3

> People tend to overlook the decay of the modern web, when in fact these numbers are extraordinary—they represent a comprehensive breakdown in the chain of custody for facts. This is a particularly good quote to sum up the article. The internet is not a repository of facts, it is a repository of facts, spam, junk, and things . Moreover, it is not the only repository of these. Link rot happens. Content is subject to…

It's depressing. I know somebody who started a business that was successful for a while and then failed. Spammers got control of the domain and now it is full of ads for a dangerous diet drug. What makes my blood boil is that it impugns the integrity of the founder who is a decent person who has nothing to do with that scam.

Domains, if not outright sold by the owner, should die then.

Re: Link rot and content drift are endemic to the web

#12

yes. i am down to HN and Reddit. Google calendar and telegram. I don't even know how to find cool stuff online anymore. Google SERPs are all business driven now unless you're research news.

Google still works, you just need to be very specific on your search criteria (e.g. site:). I do agree on the premise that good quality data, art and content has been lost to the winds of time.

Re: Link rot and content drift are endemic to the web

#13

Earlier quoted context omitted.

It's depressing. I know somebody who started a business that was successful for a while and then failed. Spammers got control of the domain and now it is full of ads for a dangerous diet drug. What makes my blood boil is that it impugns the integrity of the founder who is a decent person who has nothing to do with that scam.

Domains, if not outright sold by the owner, should die then.

And never be possible to register ever again? I feel like most easy-to-type domains would have been permanently expended in the early days of the web.

Re: Link rot and content drift are endemic to the web

#15
I expected the article to deal more with the rotting landscape of the internet, the rotting of our choices of content, the poor selection of links on the 1st page of any search engine.... I am less concerned with the fact that a link breaks than I am with what it says about the content that was there, is no longer there.

Re: Link rot and content drift are endemic to the web

#16
By the way, the technical side of this is very interesting. If you look at the tools mentioned (the wayback machine, but also perma.cc and other archival solutions), almost all of them rely on a single semi-modern tech stack that produces WARCs (web archives - ISO - ISO 28500:2017 https://iipc.github.io/warc-specifications/specifications/wa...).

The main crawler still seems to be heritrix3 (https://github.com/internetarchive/heritrix3), but there's a great little ecosystem with tools such as webrecorder and warcprox.

Still, I've read through the code of these tools and am feeling that they are failing in the face of the modern web with single page apps, mobile phone apps and walled gardens. Even newer iterations with browser automation are getting increasingly throttled and blocked and excluded from walled gardens.

Perhaps the time has come for a coordinated, decentralized but omnipresent approach to archival.

Re: Link rot and content drift are endemic to the web

#17
Knowing that it decays is what prompts us to try and save the bits worth saving.

I don't sit in the camp that everything digital must be preserved and that it's a disaster if it isn't. I try not to fight entropy in it's many manifestations. It's a shame when content disappears but I think it's also healthy to just accept it. We tend to only frame information disappearing in a negative light because we can always imagine a scenario where that information could have been valuable to someone, and that is a valid concern, I just don't think it's helpful to view it as the internet going into some downward rotting spiral and therefore every single 0 and 1 must be preserved.

The major problems of the internet seem to be almost entirely cultural currently.

Re: Link rot and content drift are endemic to the web

#18

yes. i am down to HN and Reddit. Google calendar and telegram. I don't even know how to find cool stuff online anymore. Google SERPs are all business driven now unless you're research news.

I wish there was something like Reddit that could be organized by topic, but had the simplicity of HN's design instead of the monstrocity Reddit has become. My guess is it would still succumb to the Reddit Hive Mind effect without a reasonably benevolent moderation team though. For all the times I've said that HN basically does the same thing, I have to admit that it is much better about keeping it in check.

I don’t know about userbase but old.reddit with disabled subreddit CSS is quite usable for me.

Re: Link rot and content drift are endemic to the web

#20

By the way, the technical side of this is very interesting. If you look at the tools mentioned (the wayback machine, but also perma.cc and other archival solutions), almost all of them rely on a single semi-modern tech stack that produces WARCs (web archives - ISO - ISO 28500:2017 https://iipc.github.io/warc-specifications/specifications/wa... ). The main crawler still seems to be heritrix3 ( https://github.com/inter…

> increasingly throttled and blocked and excluded from walled garden

I keep thinking back to Jacob Applebaum's stance of "facebook and the other walled gardens are the real dark web."

Post reply on HN