> People tend to overlook the decay of the modern web, when in fact these numbers are extraordinary—they represent a comprehensive breakdown in the chain of custody for facts. This is a particularly good quote to sum up the article. The internet is not a repository of facts, it is a repository of facts, spam, junk, and things . Moreover, it is not the only repository of these. Link rot happens. Content is subject to…
It's depressing. I know somebody who started a business that was successful for a while and then failed. Spammers got control of the domain and now it is full of ads for a dangerous diet drug. What makes my blood boil is that it impugns the integrity of the founder who is a decent person who has nothing to do with that scam.
Link rot and content drift are endemic to the web
11–20 of 220 posts
Re: Link rot and content drift are endemic to the web
#12yes. i am down to HN and Reddit. Google calendar and telegram. I don't even know how to find cool stuff online anymore. Google SERPs are all business driven now unless you're research news.
Re: Link rot and content drift are endemic to the web
#13Earlier quoted context omitted.
It's depressing. I know somebody who started a business that was successful for a while and then failed. Spammers got control of the domain and now it is full of ads for a dangerous diet drug. What makes my blood boil is that it impugns the integrity of the founder who is a decent person who has nothing to do with that scam.
Domains, if not outright sold by the owner, should die then.
Re: Link rot and content drift are endemic to the web
#14The second law of thermodynamics applies to the Internet as well.
Re: Link rot and content drift are endemic to the web
#15Re: Link rot and content drift are endemic to the web
#16The main crawler still seems to be heritrix3 (https://github.com/internetarchive/heritrix3), but there's a great little ecosystem with tools such as webrecorder and warcprox.
Still, I've read through the code of these tools and am feeling that they are failing in the face of the modern web with single page apps, mobile phone apps and walled gardens. Even newer iterations with browser automation are getting increasingly throttled and blocked and excluded from walled gardens.
Perhaps the time has come for a coordinated, decentralized but omnipresent approach to archival.
Re: Link rot and content drift are endemic to the web
#17I don't sit in the camp that everything digital must be preserved and that it's a disaster if it isn't. I try not to fight entropy in it's many manifestations. It's a shame when content disappears but I think it's also healthy to just accept it. We tend to only frame information disappearing in a negative light because we can always imagine a scenario where that information could have been valuable to someone, and that is a valid concern, I just don't think it's helpful to view it as the internet going into some downward rotting spiral and therefore every single 0 and 1 must be preserved.
The major problems of the internet seem to be almost entirely cultural currently.
Re: Link rot and content drift are endemic to the web
#18yes. i am down to HN and Reddit. Google calendar and telegram. I don't even know how to find cool stuff online anymore. Google SERPs are all business driven now unless you're research news.
I wish there was something like Reddit that could be organized by topic, but had the simplicity of HN's design instead of the monstrocity Reddit has become. My guess is it would still succumb to the Reddit Hive Mind effect without a reasonably benevolent moderation team though. For all the times I've said that HN basically does the same thing, I have to admit that it is much better about keeping it in check.
Re: Link rot and content drift are endemic to the web
#19Re: Link rot and content drift are endemic to the web
#20By the way, the technical side of this is very interesting. If you look at the tools mentioned (the wayback machine, but also perma.cc and other archival solutions), almost all of them rely on a single semi-modern tech stack that produces WARCs (web archives - ISO - ISO 28500:2017 https://iipc.github.io/warc-specifications/specifications/wa... ). The main crawler still seems to be heritrix3 ( https://github.com/inter…
I keep thinking back to Jacob Applebaum's stance of "facebook and the other walled gardens are the real dark web."