Live data from Hacker News

Link rot and content drift are endemic to the web

theatlantic.com

71–80 of 220 posts

Re: Link rot and content drift are endemic to the web

#71

It's an Atlantic article so it's long, and several of the comments here show that people aren't actually reading the whole thing... but I did and it's worth the time. It's not only about links being dead, it's about the lack of transparency & audit when content is changed via takedown requests, it's about dead links showing up in decades-old supreme court decisions, it's about private industry's lack of incentive for…

Having read the whole thing, it seems partially good, martially misguided, and partially terrible. The overall bent is a hand-wringing about link rot, which I thought we mostly got over a decade ago. The Internet is fundamentally ephemeral. If you see something you like, save it so you can repost it later. If you rely on someone else to keep in up indefinitely, you're being foolish. Around the edges of that main disc…

I'm not 100% convinced by your assertion that censorship only favors the rich and powerful. It can and often does, but it can also help people without power or society and large. For instance, taking down a dox for a niche YouTuber is clearly not helping a powerful person, but it's still arguably censorship.

The misinformation area is somewhat stickier, but here's a decent example: if somebody decided to hurt you by spreading rumors (let's say that you watch CP) and spends time and money to get that rumor top of any SEO and forum thread, what's the right course of action? How good are you going to feel about using speech to counteract that when the result is a Google search giving your denial in spot 1 and the accusation in spot 2?

We all have to grapple with power and the ability to abuse it, but I don't think it's effective to say power is fundamentally wrong. The conversation is more nuanced than that and has to be viewed as systems with checks on power, which means specific design-thinking.

Re: Link rot and content drift are endemic to the web

#72
I actually implemented a rule for my website: anytime I write anything and cite a link, I always also include the internet archive url as well just in case. If it's not been archived yet I submit it to be.

as an example:

"You don't have to trust me on this one, here's an article with [a bunch of data] | [*Archive link in case of link rot]"

from: https://kolemcrae.com/notebook/virtue.html

It's not perfect, but it helps reduce some of the issue.

Other than that solutions are incredibly hard to come by - you need institutions to preserve urls - through tech changes and the like, when they have very little incentive to do so. Eg. making sure they implement a redirect from the http to https sounds simple enough, but not everyone did it. Also if they switch CMSs and the like.

Re: Link rot and content drift are endemic to the web

#74
post #3

> People tend to overlook the decay of the modern web, when in fact these numbers are extraordinary—they represent a comprehensive breakdown in the chain of custody for facts. This is a particularly good quote to sum up the article. The internet is not a repository of facts, it is a repository of facts, spam, junk, and things . Moreover, it is not the only repository of these. Link rot happens. Content is subject to…

Exactly. The Atlantic author seems to be laboring under the misguided assumption that the web is somehow the same sort of thing as a library of books. Even libraries often have some degree of garbage information in them, and represent a survival story: the vast majority of books ever written are no longer in print, or even discoverable anywhere. Good stuff should be preserved, but it's not the Internet's job to someh…

Somehow I'm replying to you twice, but this time I agree and wanted to note that libraries are curated spaces as well.

Re: Link rot and content drift are endemic to the web

#75
I feel that the Internet as an archive isn't really feasible. At best, it can augment existing archival efforts such as public libraries. The fact people keep pushing off to webhosting what should be put into a library is a grave misunderstanding of the use cases for the Internet.

Re: Link rot and content drift are endemic to the web

#76

Buddhists chuckle at the notion of permanence and go back to constructing sand mandalas

Indeed yes, however much we may dislike it, change is the only constant. Of course that doesn't mean we shouldn't bother archiving, but there is no need to fret over saving every byte out there on the web.

Entropy is king. Eventually all information loses its coherency.

Re: Link rot and content drift are endemic to the web

#77

I am increasingly worried about the valuable content on YouTube. There are so many old live concerts, useful how-to videos and other cultural treasures amidst all the junk. I suspect that one day, they will make their ads unblockable by embedding them in the video files. I sure hope that some people are downloading the valuable stuff and stashing it away to load onto YouTube's successor.

It’s really up to you who cares about something to archive it. I managed to find a torrent of early days video games from my region that has almost been lost to time. Luckily I found a discord and could coax someone to hop on to seed it. If I had waited 10 more years they might have been gone for good.

Everyone’s assuming that data now stays on the internet forever because it’s so massive. It’s usually one or two people who keep the flame alive

Re: Link rot and content drift are endemic to the web

#78
post #63

It's amazing how little some trusted institutions care about this. For example, the BBC has been bragging about how many people rely on their coverage of the pandemic, but have an obnoxious habit of repeatedly overwriting old articles with new ones on similar topics and not keeping the old versions available. The history of a once-in-a-century pandemic with huge local and global impacts is literally being overwritten…

> so unless historians dig deep in third-party archives, they'd never understand where that belief came from

I expect future historical tooling will exist to solve exactly this problem. Assuming Archive.org and the like nabbed it, the evidence is all there for future generations to see.

Re: Link rot and content drift are endemic to the web

#79

Earlier quoted context omitted.

I used to be a data hoarder but I learned to let it go. I save the important stuff, just like the Buddhist monks do. How much of the internet is really worth saving? What will Geocities mean to anyone 50 years from now? The internet is a dynamic process evolving in real time. No man ever steps in the same river twice, for it's not the same river and he's not the same man Heraclitus

I am a fan of that quote and I now personally eschew clutter. However I'm referring to the examples in the article such as the supreme court justice referencing links that no longer existed. Paraphrasing the article, >75% of links from the 90s are defunct. Sure Geocities may not have value to many, but an astonishing number of links in court rulings and law documents are leading to dead ends. I can see how this could…

Google's recent invention of text links should help this:

https://en.wikipedia.org/wiki/Filler_text#:~:text=%22Now%20i....

# signifies an anchor

:~:text= signifies a text link

%22Now%20is%20the%20time%20for,21%20(1918). says show me the text between "Now is the time for" and "21 (1918)."

Re: Link rot and content drift are endemic to the web

#80
post #3

> People tend to overlook the decay of the modern web, when in fact these numbers are extraordinary—they represent a comprehensive breakdown in the chain of custody for facts. This is a particularly good quote to sum up the article. The internet is not a repository of facts, it is a repository of facts, spam, junk, and things . Moreover, it is not the only repository of these. Link rot happens. Content is subject to…

It's depressing. I know somebody who started a business that was successful for a while and then failed. Spammers got control of the domain and now it is full of ads for a dangerous diet drug. What makes my blood boil is that it impugns the integrity of the founder who is a decent person who has nothing to do with that scam.

Nodes in keywordspace don’t die with the businesses that created them. The popularity of domains and links and words and phrases are permanently altered by the existence of the business. It’s a digital footprint like how a real-world business leaves a physical footprint. Some footprints are harmless - just a memory of activity that once happened. Other footprints cause lasting harm, like contaminated soil.

Abandoned formerly-popular domains create a kind of long-tail info-environmental impact, just like an abandoned warehouse can become a real-world hazard.

Maybe we need a digital superfund process.

Post reply on HN