Earlier quoted context omitted.
The BBC publishes nothing but garbage and there are other extant sources that are more durable. It's fine to forget things. We're missing entire libraries of classical literature from great authors which would be nice to have. Missing documentary sewage isn't a tragedy. This should show us that most of the web isn't worth preserving anyway, much like McDonald's burger wrappers aren't worth preserving like sacred arti…
I remember how calling the BBC garbage a few years ago got your comment heavily downvoted here. They'd tell you that they were the best thing since sliced bread and that they were good because both the left and the right hated them, as if that meant something. Now it seems everybody is recognising the BBC for what they are: utter shite.
Link rot and content drift are endemic to the web
121–130 of 220 posts
Re: Link rot and content drift are endemic to the web
#122How they get implemented in solving this problem is the question
Re: Link rot and content drift are endemic to the web
#123I am increasingly worried about the valuable content on YouTube. There are so many old live concerts, useful how-to videos and other cultural treasures amidst all the junk. I suspect that one day, they will make their ads unblockable by embedding them in the video files. I sure hope that some people are downloading the valuable stuff and stashing it away to load onto YouTube's successor.
you're worried about....the ads? Having ads around doesn't make the content any less valuable. We're talking about the content itself still being available, who cares if there's some ads keeping the system up if all those live concerts and how-to vids are preserved forever
Re: Link rot and content drift are endemic to the web
#124Earlier quoted context omitted.
On the bright side, using a tool like Internet archive it should be easy to filter out which articles were removed and/or edited by the BBC, in a way highlighting the most historically important articles.
I mean yes, if you are extremely motivated. But the wayback machine is pretty klunky and slow, honestly. And there's no good "diff" view that summarizes the changes to a URL over time AFAIK.
This is very true. Sometimes it takes 5-10 seconds to load the calendar view for an archived page, and another 5-10, or more, to load a snapshot.
They have a ton of data to manage with limited resources, but it still seems it should be possible to go faster than this. If there's just not enough budget for I/O, maybe they could offer a donate-for-data-dump option, where you can donate in exchange for loading data of interest (say, BBC archives) into a storage medium or query engine of one's choice, so one could do research at a much faster pace.
Re: Link rot and content drift are endemic to the web
#125Earlier quoted context omitted.
WARC can record and replay single-page apps, but it struggles with knowing where a "page" begins and ends. There was a time when I was furious with the web going to hell and I investigated the possibility of "web without browsers" that started with making a WARC capture of page and putting pages through extensive filtering and classification before the user sees anything. With interactive capturing you can push a but…
I'm completely out of the loop on something like this, but could you in theory apply some kind of ML to identify the end of pages to assist with good page captures?
At the time I was most bothered by the slow load times of web pages and blaming this phenomenon:
https://www.sjsu.edu/faculty/watkins/samplemax4.htm
particularly that if you take the max of N random variables, the expectation value you get gets worse as N increases -- that is, the page isn't done loading until the slowest http request completes.
So I saw the "knowing when the page is done" problem as being particularly core, and it would be if the goal was to "win the race" against a conventional web browser.
If you were (say) preloading all the links submitted to hacker news you might be able to tolerate the system taking 5 minutes to process an incoming page. (See archive.is)
Today I've noticed that sites like Wired are giving up on complaining about my anti-track and ad-blocker and they just load the page partially which would drive me crazy if I was serious about debugging.
Re: Link rot and content drift are endemic to the web
#126Re: Link rot and content drift are endemic to the web
#127Earlier quoted context omitted.
I remember how calling the BBC garbage a few years ago got your comment heavily downvoted here. They'd tell you that they were the best thing since sliced bread and that they were good because both the left and the right hated them, as if that meant something. Now it seems everybody is recognising the BBC for what they are: utter shite.
I see this as a more general pattern on HN: Opinions not-yet-adopted by academia are often downvoted instead of being argued with. This stifles innovation because alternative opinions do not even show up in the casual reader's screen.
Better to post links to trusted sources and let people judge for themselves.
Re: Link rot and content drift are endemic to the web
#128Earlier quoted context omitted.
The BBC publishes nothing but garbage and there are other extant sources that are more durable. It's fine to forget things. We're missing entire libraries of classical literature from great authors which would be nice to have. Missing documentary sewage isn't a tragedy. This should show us that most of the web isn't worth preserving anyway, much like McDonald's burger wrappers aren't worth preserving like sacred arti…
Reading one of the BBC's technical articles, a cyber security news item, they had 3 errors in the first paragraph. I didn't bother reading to the end of the article. I'm glad I no longer pay for a TV license.
You notice it cos you are in tech, but the same happens in financial news, science and even sport. Go read a tech publication.
For shits and giggle I did once try to get a technical story on how to copy DVD's published - it got very heavily edited! http://news.bbc.co.uk/2/hi/science/nature/1987665.stm
(I'm a former + early BBC News website employee)
Re: Link rot and content drift are endemic to the web
#129Buddhists chuckle at the notion of permanence and go back to constructing sand mandalas
The invention of wood block printing 1300 years ago in the Tang was apparently specifically motivated by the desire to preserve and reproduce Buddhist sutras; the oldest surviving documents printed with movable type, from 900 years ago, are also Buddhist texts.
Of course the Tripitaka is not permanent; it will be lost some day. But you seem to be implicitly claiming that Buddhists do not apply effort to preserving information and in particular textual records, because they know that ultimately they will be lost. In fact, the truth is quite the opposite, and believing your implicit claim would require almost complete ignorance of Buddhism, printing technology, and South Asian classical studies.