Live data from Hacker News

Link rot and content drift are endemic to the web

theatlantic.com

151–160 of 220 posts

Re: Link rot and content drift are endemic to the web

#151
I doubt people will ever care about this enough for it to have momentum, but there are well known technological solutions: content addressable file storage. If you do that the url is always tied to the file content itself. Of course this requires documents to actually be documents. So I don't think it works for any modern business model.

Re: Link rot and content drift are endemic to the web

#152

I am increasingly worried about the valuable content on YouTube. There are so many old live concerts, useful how-to videos and other cultural treasures amidst all the junk. I suspect that one day, they will make their ads unblockable by embedding them in the video files. I sure hope that some people are downloading the valuable stuff and stashing it away to load onto YouTube's successor.

> that I should just pay a subscription fee to avoid their ad crap

Fuck no. Do not do this. Use youtube-dl[0], and maintain local copies of anything useful you can find.

0: http://youtube-dl.org/

Re: Link rot and content drift are endemic to the web

#153

I am increasingly worried about the valuable content on YouTube. There are so many old live concerts, useful how-to videos and other cultural treasures amidst all the junk. I suspect that one day, they will make their ads unblockable by embedding them in the video files. I sure hope that some people are downloading the valuable stuff and stashing it away to load onto YouTube's successor.

How many how to videos, memes, concerts, etc, are really that important? In fifty years how many people will care? How many people should care because it would mean ignoring the huge volume of newer stuff? A hundred years? Two hundred? I haven't even read or seen many of the existing cultural artifacts we have from past decades and centuries, what would I do if orders of magnitudes more of them had been preserved? In…

These subjective questions are pointless to try and answer.

The idea is that a future person could freely deep dive through a rich well indexed history of media about whatever specifically interests them

I wish people would stop trying to assess the value of a given piece of media and just tag and archive the stuff.

For instance, high quality footage of live music from 100 years ago would be very interesting to some.

Re: Link rot and content drift are endemic to the web

#154

I actually implemented a rule for my website: anytime I write anything and cite a link, I always also include the internet archive url as well just in case. If it's not been archived yet I submit it to be. as an example: "You don't have to trust me on this one, here's an article with [a bunch of data] | [*Archive link in case of link rot]" from: https://kolemcrae.com/notebook/virtue.html It's not perfect, but it help…

> a rule for my website

Note that you should also have a rule to save the link content locally, to avoid single-point-of-failure problems in the unlikely-but-catastrophic case that archive.org itself goes down. (Cf the attempts to attack them over their National Emergency Library programme last year.)

Re: Link rot and content drift are endemic to the web

#155
post #128

Earlier quoted context omitted.

Reading one of the BBC's technical articles, a cyber security news item, they had 3 errors in the first paragraph. I didn't bother reading to the end of the article. I'm glad I no longer pay for a TV license.

The BBC (News's) tech section isn't aimed at you. Inaccuracies shouldn't be there but often they will dumb down or gloss over stuff for the mainstream audience they are aiming at. You notice it cos you are in tech, but the same happens in financial news, science and even sport. Go read a tech publication. For shits and giggle I did once try to get a technical story on how to copy DVD's published - it got very heavily…

Aside: That old version of BBC News is an absolute gem of history. Especially looking at some of the recommended sidebar stories:

> Britons 'baffled over euro rate'

> Wireless internet arrives in China

> Mobile spam on the rise

Fascinating to see how much our problems have stayed the same, despite the changing context.

I hope this is considered 'archived' and not 'forgotten'.

Re: Link rot and content drift are endemic to the web

#156

Earlier quoted context omitted.

I download everything I like. Storage is cheap now. Plus smaller sites will start disappearing because of regulatory capture. It will not be possible to run a forum or similar site in few years.

We say that a lot, but it's not that cheap if you're using redundancy and backups. It's cheap if you don't care too much about the data.

Cataloging and indexing and searchability are also not cheap if you are doing all of that on your own time.

Re: Link rot and content drift are endemic to the web

#157
It's not a technology problem, it's an incentive problem.

Had the web somehow been centralized (I have no idea what that would even look like), content still would not be archived, it would be constantly changed, and subject to censorship. Just like in a decentralized web, perhaps even more so.

Archiving costs lots of money (and costs keep growing if you only add and never take away), can be highly challenging (in the case of web apps or complex dependencies), whilst providing zero immediate reward for the organization carrying this heavy load. Not only is there no incentive, many couldn't even afford to if they wanted to.

And it gets worse still. Digital archiving means paying forever. Imagine paying a 100 years of electricity, hardware replacements, migrations. The entity (business, person) is long gone before that.

As a ridiculous example of this: Facebook has several very large idle content data centers. Mega scale buildings full of servers storing photos of Facebook users they haven't accessed in years, and likely never will again. Yet should a user do this, they expect the photo to still be there.

That's why I believe the problem should be addressed with more pragmatism. Focus on things of unquestionable long term value, and think of a good solution for this smaller scope.

Re: Link rot and content drift are endemic to the web

#158
post #71

Earlier quoted context omitted.

I'm not 100% convinced by your assertion that censorship only favors the rich and powerful. It can and often does, but it can also help people without power or society and large. For instance, taking down a dox for a niche YouTuber is clearly not helping a powerful person, but it's still arguably censorship. The misinformation area is somewhat stickier, but here's a decent example: if somebody decided to hurt you by…

> if somebody decided to hurt you by spreading rumors We have libel laws to address that. GPs point is that Google, Facebook, et. al. are premptively censoring non-mainstream content just to protect themselves. They don't really care about the public.

> just to protect themselves

To protect themselves from the public. Whether it's because consumers might take their business (and their data) elsewhere in disgust at what a particular platform is turning into, or because democratically elected lawmakers could start imposing sanctions or new regulations.

Companies are always looking out for themselves, that's a given. But that doesn't mean their actions are completely divorced from public opinion.

Re: Link rot and content drift are endemic to the web

#159
For anyone that wants to help with this, check out the Archive Team Warrior project. You can donate bandwidth and some CPU cycles to archiving different parts of the web. There's a VM image you can download that makes it really easy.

https://wiki.archiveteam.org/index.php/ArchiveTeam_Warrior

You can choose to help archive reddit, pastebin, URL shorteners and other ephemeral parts of the internet

https://wiki.archiveteam.org/index.php/Warrior_projects

I've also taken to updating the citations in Wikipedia articles with archive links.

Re: Link rot and content drift are endemic to the web

#160
post #69

Earlier quoted context omitted.

On the bright side, using a tool like Internet archive it should be easy to filter out which articles were removed and/or edited by the BBC, in a way highlighting the most historically important articles.

I mean yes, if you are extremely motivated. But the wayback machine is pretty klunky and slow, honestly. And there's no good "diff" view that summarizes the changes to a URL over time AFAIK.

Slow, yes. Klunky? Absolutely not. When you make a request to the Internet Archive, you're searching through a massive amount of data. The fact that it only takes a few seconds to pull up a decade-old webpage is amazing.
Post reply on HN