Live data from Hacker News

Link rot and content drift are endemic to the web

theatlantic.com

131–140 of 220 posts

Re: Link rot and content drift are endemic to the web

#131

I am increasingly worried about the valuable content on YouTube. There are so many old live concerts, useful how-to videos and other cultural treasures amidst all the junk. I suspect that one day, they will make their ads unblockable by embedding them in the video files. I sure hope that some people are downloading the valuable stuff and stashing it away to load onto YouTube's successor.

How many how to videos, memes, concerts, etc, are really that important? In fifty years how many people will care? How many people should care because it would mean ignoring the huge volume of newer stuff? A hundred years? Two hundred? I haven't even read or seen many of the existing cultural artifacts we have from past decades and centuries, what would I do if orders of magnitudes more of them had been preserved? In…

>I haven't even read or seen many of the existing cultural artifacts we have from past decades and centuries, what would I do if orders of magnitudes more of them had been preserved?

You might have a better understanding of the culture that produced them. You might appreciate a work of art that would otherwise not exist. We have graffiti from Pompeii, we know Ea-nasir sold cheap copper in ancient Ur 3700-odd years ago, but we've lost countless works of literature, music and film, some by the greatest masters of their age. What artifacts of culture survive the scouring sands of time is often a matter of happenstance, rather than quality.

Chances are almost everything our species has produced culturally, scientifically and artistically - the whole corpus of our knowledge output over the last century - is going to vanish within a generation or two anyway, simply because the digital foundation into which we've transferred so much of it is brittle and ephemeral. If we want to leave anything behind for future generations at all besides climate change, pollution and nuclear waste, we should save as much as possible rather than only what we consider to be relevant.

Re: Link rot and content drift are endemic to the web

#132
post #69

Earlier quoted context omitted.

On the bright side, using a tool like Internet archive it should be easy to filter out which articles were removed and/or edited by the BBC, in a way highlighting the most historically important articles.

Wayback Machine censors many websites(like 4chan) from being 'saved'. Wayback Machine also removes previously archived videos/websites in certain cases. They are not neutral.

Fortunately 4chan has (unofficial) archives, but some content was probably lost.

Re: Link rot and content drift are endemic to the web

#133
I always thought something like Ethereum could solve this type of thing; that is, if the content itself lived inside the blockchain. Obviously for larger formats that wouldn't work, but for many text based or lower resolution image formats, it wouldn't be too much overhead to just inject it all into the blockchain.

Re: Link rot and content drift are endemic to the web

#135

By the way, the technical side of this is very interesting. If you look at the tools mentioned (the wayback machine, but also perma.cc and other archival solutions), almost all of them rely on a single semi-modern tech stack that produces WARCs (web archives - ISO - ISO 28500:2017 https://iipc.github.io/warc-specifications/specifications/wa... ). The main crawler still seems to be heritrix3 ( https://github.com/inter…

I think “right to be forgotten” is important and I’m generally against everlasting social media posts, but for copyrighted works, we really need a centralized Library of Congress that acts to archive these. In order for that to happen there needs to be an equivalent “publishing” mechanism for the web - where the user says - I created something and I want it to be archived. This would cover things that exist behind a paywall or are only delivered as newsletters.

Re: Link rot and content drift are endemic to the web

#136
post #121
post #92

Earlier quoted context omitted.

I remember how calling the BBC garbage a few years ago got your comment heavily downvoted here. They'd tell you that they were the best thing since sliced bread and that they were good because both the left and the right hated them, as if that meant something. Now it seems everybody is recognising the BBC for what they are: utter shite.

I see this as a more general pattern on HN: Opinions not-yet-adopted by academia are often downvoted instead of being argued with. This stifles innovation because alternative opinions do not even show up in the casual reader's screen.

Absent an explanation of why I've annoyed people, I get as much of a dopamine hit from downvotes as upvotes. I'd rather be polarising than forgettable.

Whenever I take an unpopular stance I remind myself of Rick Sanchez's wise words, "Your boos mean nothing, I've seen what makes you cheer".

Re: Link rot and content drift are endemic to the web

#137
This is why I'm building/curating my personal archive with stuff that I think may be worthwhile saving (not only for myself necessarily).

Perhaps there will be many personal archives like mine that one day can be shared in a similar vein to copy parties.

We will need to treat the information we find online with its impermanence in mind (as authors, making things easy to copy, and consumers, copying stuff).

Perhaps it is this mindset that, when sufficiently prevalent, could make the internet more like a library again; weed out the garbage und curate the nuggets.

Btw I think archive.org is doing God's work but I don't believe any amount of coding and crawling will be able to save everything (nor should it). It can capture some raw data for (future AI?) historians to sift through though.

Re: Link rot and content drift are endemic to the web

#139
post #69

Earlier quoted context omitted.

On the bright side, using a tool like Internet archive it should be easy to filter out which articles were removed and/or edited by the BBC, in a way highlighting the most historically important articles.

Wayback Machine censors many websites(like 4chan) from being 'saved'. Wayback Machine also removes previously archived videos/websites in certain cases. They are not neutral.

Don’t spread misinformation. The Wayback Machine is not censoring 4chan. 4chan is ‘censoring’ the Wayback Machine.

https://www.4chan.org/robots.txt

Also, I let a domain of mine expire and the new domain owner (which just plastered ads) had a robots.txt that retroactively removed my “previously archived website” from the Wayback Machine.

Re: Link rot and content drift are endemic to the web

#140
post #92

Earlier quoted context omitted.

The BBC publishes nothing but garbage and there are other extant sources that are more durable. It's fine to forget things. We're missing entire libraries of classical literature from great authors which would be nice to have. Missing documentary sewage isn't a tragedy. This should show us that most of the web isn't worth preserving anyway, much like McDonald's burger wrappers aren't worth preserving like sacred arti…

I remember how calling the BBC garbage a few years ago got your comment heavily downvoted here. They'd tell you that they were the best thing since sliced bread and that they were good because both the left and the right hated them, as if that meant something. Now it seems everybody is recognising the BBC for what they are: utter shite.

A few years ago, any comment that didn't add new insight to a topic would get downvoted. I remember once reading a comment where the response was a quip, and someone replied "this response was funny but we don't want this site to become Reddit so I downvoted you".
Post reply on HN