Live data from Hacker News

Link rot and content drift are endemic to the web

theatlantic.com

121–130 of 220 posts

Re: Link rot and content drift are endemic to the web

#121
post #92

Earlier quoted context omitted.

The BBC publishes nothing but garbage and there are other extant sources that are more durable. It's fine to forget things. We're missing entire libraries of classical literature from great authors which would be nice to have. Missing documentary sewage isn't a tragedy. This should show us that most of the web isn't worth preserving anyway, much like McDonald's burger wrappers aren't worth preserving like sacred arti…

I remember how calling the BBC garbage a few years ago got your comment heavily downvoted here. They'd tell you that they were the best thing since sliced bread and that they were good because both the left and the right hated them, as if that meant something. Now it seems everybody is recognising the BBC for what they are: utter shite.

I see this as a more general pattern on HN: Opinions not-yet-adopted by academia are often downvoted instead of being argued with. This stifles innovation because alternative opinions do not even show up in the casual reader's screen.

Re: Link rot and content drift are endemic to the web

#123

I am increasingly worried about the valuable content on YouTube. There are so many old live concerts, useful how-to videos and other cultural treasures amidst all the junk. I suspect that one day, they will make their ads unblockable by embedding them in the video files. I sure hope that some people are downloading the valuable stuff and stashing it away to load onto YouTube's successor.

you're worried about....the ads? Having ads around doesn't make the content any less valuable. We're talking about the content itself still being available, who cares if there's some ads keeping the system up if all those live concerts and how-to vids are preserved forever

The ads make content less valuable. They distract and mislead and manipulate emotions, especially when the youtuber sponsors something during the video itself. The content isn’t abstract and isolated, the content and the ads come as a bundle. The digital procedural product placement on the way will make this 10x worse

Re: Link rot and content drift are endemic to the web

#124
post #69

Earlier quoted context omitted.

On the bright side, using a tool like Internet archive it should be easy to filter out which articles were removed and/or edited by the BBC, in a way highlighting the most historically important articles.

I mean yes, if you are extremely motivated. But the wayback machine is pretty klunky and slow, honestly. And there's no good "diff" view that summarizes the changes to a URL over time AFAIK.

But the wayback machine is pretty klunky and slow, honestly.

This is very true. Sometimes it takes 5-10 seconds to load the calendar view for an archived page, and another 5-10, or more, to load a snapshot.

They have a ton of data to manage with limited resources, but it still seems it should be possible to go faster than this. If there's just not enough budget for I/O, maybe they could offer a donate-for-data-dump option, where you can donate in exchange for loading data of interest (say, BBC archives) into a storage medium or query engine of one's choice, so one could do research at a much faster pace.

Re: Link rot and content drift are endemic to the web

#125
post #106

Earlier quoted context omitted.

WARC can record and replay single-page apps, but it struggles with knowing where a "page" begins and ends. There was a time when I was furious with the web going to hell and I investigated the possibility of "web without browsers" that started with making a WARC capture of page and putting pages through extensive filtering and classification before the user sees anything. With interactive capturing you can push a but…

I'm completely out of the loop on something like this, but could you in theory apply some kind of ML to identify the end of pages to assist with good page captures?

Probably. Certainly the more you spent on it the better you could do.

At the time I was most bothered by the slow load times of web pages and blaming this phenomenon:

https://www.sjsu.edu/faculty/watkins/samplemax4.htm

particularly that if you take the max of N random variables, the expectation value you get gets worse as N increases -- that is, the page isn't done loading until the slowest http request completes.

So I saw the "knowing when the page is done" problem as being particularly core, and it would be if the goal was to "win the race" against a conventional web browser.

If you were (say) preloading all the links submitted to hacker news you might be able to tolerate the system taking 5 minutes to process an incoming page. (See archive.is)

Today I've noticed that sites like Wired are giving up on complaining about my anti-track and ad-blocker and they just load the page partially which would drive me crazy if I was serious about debugging.

Re: Link rot and content drift are endemic to the web

#126
I think out obsession with copyright and attribution of ideas isn't helping. Many times you'll see people reference a page or PDF, which long ago became a broken link. Not one person bothers to paraphrase or copy relevant sections from it. And the wayback machine can't cover everything.

Re: Link rot and content drift are endemic to the web

#127
post #121
post #92

Earlier quoted context omitted.

I remember how calling the BBC garbage a few years ago got your comment heavily downvoted here. They'd tell you that they were the best thing since sliced bread and that they were good because both the left and the right hated them, as if that meant something. Now it seems everybody is recognising the BBC for what they are: utter shite.

I see this as a more general pattern on HN: Opinions not-yet-adopted by academia are often downvoted instead of being argued with. This stifles innovation because alternative opinions do not even show up in the casual reader's screen.

“Someone said it on Hacker News” carries no weight. Why should anyone take our comments seriously if they don’t recognize the username? I don’t see this as a bug.

Better to post links to trusted sources and let people judge for themselves.

Re: Link rot and content drift are endemic to the web

#128

Earlier quoted context omitted.

The BBC publishes nothing but garbage and there are other extant sources that are more durable. It's fine to forget things. We're missing entire libraries of classical literature from great authors which would be nice to have. Missing documentary sewage isn't a tragedy. This should show us that most of the web isn't worth preserving anyway, much like McDonald's burger wrappers aren't worth preserving like sacred arti…

Reading one of the BBC's technical articles, a cyber security news item, they had 3 errors in the first paragraph. I didn't bother reading to the end of the article. I'm glad I no longer pay for a TV license.

The BBC (News's) tech section isn't aimed at you. Inaccuracies shouldn't be there but often they will dumb down or gloss over stuff for the mainstream audience they are aiming at.

You notice it cos you are in tech, but the same happens in financial news, science and even sport. Go read a tech publication.

For shits and giggle I did once try to get a technical story on how to copy DVD's published - it got very heavily edited! http://news.bbc.co.uk/2/hi/science/nature/1987665.stm

(I'm a former + early BBC News website employee)

Re: Link rot and content drift are endemic to the web

#129

Buddhists chuckle at the notion of permanence and go back to constructing sand mandalas

Buddhists have preserved most of the Tripitaka for 2500 years, and for the first 500 years it was memorized and transmitted orally from generation to generation. Buddhist monks today spend significant amounts of their time memorizing parts of it. Printed, it's about 12000 pages; it's been translated into many languages, but not all of it has been translated into English yet. Thanissaro Bhikkhu has been working on it for 20 years, publishing his translations under CC-BY, and may finish the job before he dies. Aside from its value to devotees, the Tripitaka is one of our best historical sources about everyday life in South Asia 2500 years ago.

The invention of wood block printing 1300 years ago in the Tang was apparently specifically motivated by the desire to preserve and reproduce Buddhist sutras; the oldest surviving documents printed with movable type, from 900 years ago, are also Buddhist texts.

Of course the Tripitaka is not permanent; it will be lost some day. But you seem to be implicitly claiming that Buddhists do not apply effort to preserving information and in particular textual records, because they know that ultimately they will be lost. In fact, the truth is quite the opposite, and believing your implicit claim would require almost complete ignorance of Buddhism, printing technology, and South Asian classical studies.

Post reply on HN