I have 10 years of lovingly curated YouTube videos playlists, which, now when I look into the older ones, are a barren wasteland of "Video removed" or "Video not available". It is heartbreaking. Is there any way I can prevent this from happening?
Internet Data Is Rotting
91–100 of 116 posts
Re: Internet Data Is Rotting
#92I actually have a different problem -- not sure it is one that I can legally solve. I have 10 years of lovingly curated YouTube videos playlists, which, now when I look into the older ones, are a barren wasteland of "Video removed" or "Video not available". It is heartbreaking. Is there any way I can prevent this from happening?
Re: Internet Data Is Rotting
#93This needs to be solved on the protocol level. Of course, the players who have control over our protocols are exactly the people who don't want this to be solved at all. The next best thing would be to redefine what "bookmarking" is. When I bookmark a page, I want it to be permanently stored on my local machine and full-text indexed. In fact, it's rather ridiculous that after 25 years browsers don't have anything of…
Best use of time and money is probably http://pinboard.in/ ? Seems like someone I can trust not to sell out and shut down.
Re: Internet Data Is Rotting
#94Earlier quoted context omitted.
I think if the last 10 years have taught us anything, it's that preserving the past does nothing to impede the changing moral zeitgeist. More records simply mean more people to attack for holding an opinion that has simply gone out of fashion. If the past decade wasn't characterized by tribalism and moral hysteria I'd be more inclined to worry about stringent preservation. At this point, I'm not really comfortable wi…
Embrace humanity. It's not that nice, but pretending otherwise only puts you at a disadvantage for understanding and dealing with the world. Our history is an inseparable part of who we are, and it should not be forgotten.
Re: Internet Data Is Rotting
#95Yeah it’s annoying that links get broken. But maybe it’s better this way. There’s something about modern tech that has turned all of us into digital horders. I (we?) have backups and backups of backups and redundant RAID servers with every version of every file so that no byte shall ever perish. I have essays still that I wrote in highschool nearly 20 years ago. To what end? I’m partial to material minimalism. Why no…
My common response is that there are real world costs to having too much stuff, data is pretty small to keep around. You can fit a RAID 5 box with 30 TB of space in a shoebox for $15/month, that is enough to keep pretty much any content that you ever consume so that if you want anything it still exists. My parents and grandparents hoarded files and documents to no end and I have been going through some of them and th…
I'm glad that data is rotting on the Internet. In fact, I'll go one step farther and say that data should rot more quickly on the Internet. There's no reason that some scummy marketing algorithm should have access to my high-school social media posts. There's no reason that governments or private individuals with a grudge ought to be able to through the history of everything I've written, no matter how off-the-cuff, in order to dredge up something that makes me look like an undesirable when it's taken out of context.
If an individual wants to save a particular piece of information, that should be their choice. Otherwise, by default, information ought to disappear from the Internet.
Re: Internet Data Is Rotting
#96We should aim at browsing the Internet by date . We're moving everything there ignoring the fact that as it is there is no built-in permanence. We're accustomed to that after Gutenberg: it wasn't that easy losing every copy of an important document. Now it is, things disappear and we're getting into a drifting cultural bubble impossible to trace back. The Internet Archive is doing God's work, but it's not enough. If…
The effort has been made. Look at what happened with Tibet:
> Caroe pulled off a rather notorious subterfuge in order to buttress the British claim to Tawang: he published the Simla Convention for the first time in 1938 with a note misrepresenting that it had included settlement of the border (and alienation of Tawang); and he arranged for the publication of official Survey of India maps that, for the first time, showed the McMahon Line as the official boundary. To advance the narrative, he also corresponded with commercial atlas publishers to put the McMahon Line on their maps as well.
> In a telling indication of Caroe’s jiggery-pokery, to avoid the awkward question of why he was first publishing the Simla Convention twenty-four years after the fact in 1938, he instead arranged for the surreptitious printing of a spurious back-dated edition of Aitchison, deleting the original note about the Chinese government’s non-signature, and replacing it with a lengthy note stating, quite falsely, that “The [Simla] Convention included a definition of boundaries…”
> Since 1) the McMahon Line had been concluded in secret bilateral negotiations between Tibet and Great Britain outside the Convention and 2) the Chinese had officially refused to recognize any bilateral agreement, boundary or otherwise, between Tibet and Great Britain and 3) had declined to sign the Simla Convention itself and 4) had notified Great Britain in 1914 that the specific sticking point was “the boundaries” this was hoo-hah.
> The replacement copy was distributed to various libraries with instructions to withdraw and destroy the original edition.
> The subterfuge was only discovered in 1963 when J.A. Addis, a British diplomat, discovered a surviving copy of the original edition at Harvard and compared it to Caroe’s version.
( http://www.unz.com/plee/the-myth-of-the-mcmahon-line/ )
Wikipedia confirms this, if you look hard, in a shockingly non-judgmental way:
> Simla was initially rejected by the Government of India as incompatible with the 1907 Anglo-Russian Convention. The official treaty record, C.U. Aitchison's A Collection of Treaties, was published with a note stating that no binding agreement had been reached at Simla. Since the condition (agreement with China) specified by the accord was not met, the Tibetan government didn't agree with the McMahon Line.
> The Anglo-Russian Convention was renounced by Russia and Britain jointly in 1921, but the McMahon Line was forgotten until 1935, when interest was revived by civil service officer Olaf Caroe. The Survey of India published a map showing the McMahon Line as the official boundary in 1937. In 1938, the British published the Simla Convention in Aitchison's Treaties. A volume published earlier was recalled from libraries and replaced with a volume that includes the Simla Convention together with an editor's note stating that Tibet and Britain, but not China, accepted the agreement as binding. The replacement volume has a false 1929 publication date.
Re: Internet Data Is Rotting
#97Yeah it’s annoying that links get broken. But maybe it’s better this way. There’s something about modern tech that has turned all of us into digital horders. I (we?) have backups and backups of backups and redundant RAID servers with every version of every file so that no byte shall ever perish. I have essays still that I wrote in highschool nearly 20 years ago. To what end? I’m partial to material minimalism. Why no…
Managing what data you keep is just too time consuming. Hoarding it all is simpler.
Re: Internet Data Is Rotting
#98Earlier quoted context omitted.
My common response is that there are real world costs to having too much stuff, data is pretty small to keep around. You can fit a RAID 5 box with 30 TB of space in a shoebox for $15/month, that is enough to keep pretty much any content that you ever consume so that if you want anything it still exists. My parents and grandparents hoarded files and documents to no end and I have been going through some of them and th…
There are real-world costs to information hoarding as well, when it's done by people who are not you and whose incentives are not aligned with your own. I'm glad that data is rotting on the Internet. In fact, I'll go one step farther and say that data should rot more quickly on the Internet. There's no reason that some scummy marketing algorithm should have access to my high-school social media posts. There's no reas…
Re: Internet Data Is Rotting
#99A good history preservation should allow you to somehow "browse" it, as if the historical system is still alive. How the website worked, how it was used, that's all parts of the history. If old operating systems and programs are preserved, there is no reason not to preserve websites in this way.
Back in the old days, many systems are federated and/or distributed, which means the software and the protocol are two separate entities. You use a newsreader, which talks the NNTP protocol to obtain news from a Usenet newsgroup. If you want to preserve history, you can (a) archive the newsreader program with source code, and (b) archive all the data on the NNTP server. That's exactly what has been done already, if you load a Usenet archive to your newsreader, pretty much you would have the experience similar to how Larry Wall browsed the Usenet back in the late 80s, at worst you need to rewrite a compatible "mock" server, but that's all. On the other hand, little of the early BBS systems have been preserved, once the server is gone, everything is gone.
The transformation to the web, means now the platform (a web community) = protocol (backend database format) = user interface (HTML/CSS), they're all tightly coupled together. It creates several problems:
(1) The "internal state" cannot be archived. A website is a system with constantly updating parameters, and often they are not stored. Simple examples: (a) On Hacker News, I cannot see what was shown on the frontpage yesterday retroactively, (b) A user changes his/her avatar, now we had no idea how the old avatar used to look like, and (c) an early user has been banned from the forum, now his/her personal profile is inaccessible, (d) on some social media platforms, sometimes a old post may be raised from the dead by renewed interests (look, how stupid this comment was!), and now suddenly it's flooded by new posts, leaving no trace of how it used to look like.
(2) The "reader/user interface" cannot be archived. You must have seen something like this: You changed the website frontend, superficially, lots of "conservative" users complained, but the point is: now the old frontend and its "look-and-feel" is lost. If it was a simple CSS file, there are chances to bring it back, but if it was a major rewrite of frontend code, now history is gone forever. And in the lifetime of a website, the design and architecture is likely to be changed many times.
As a result, even if a website and all its content is still alive, it may already be a shadow of its past for a long time, don't even mention to preserve it! And currently, there are two ways to archive the web, both are flawed:
(1) Preserve the HTML at the surface. It's good for single pages, but you cannot browse a website in this way at all. None of the button on the website would work.
(2) Preserve the database. For example, using the API to save posts, or dumping the database - the frontend and reader are not preserved. Using Hacker News as an example, now every single post is archived, but it's far from a full experience, at least you should be able to click someone's username and see all the posts.
Now more and more websites are powered by JavaScript, makes the problem even worse. You are now literally running a program on your computer without any control over it. Once the platform is gone, no archive can save you.
What is the solution? I guess there's no full solution, but there are some possibilities:
(1) Wikipedia-like websites already have builtin version control, but it's very difficult to browse the historical version of the entire website. Systems like this can improve the frontend / user interface to allow a user to "lock on" a historical date.
(2) When building an all-Javascript website, spend some energy to build a plain HTML version as well, it may help avoiding the upcoming digital dark age.
(3) If you are going to close a website, perhaps it may be a good idea to make your internal backups of database and codebase at different years publicly available with sensitive information removed, and allow everyone to setup and run a replicated version. It's infeasible for a big website, but it may be a workable idea for a small community.
And I can imagine the archeologist from the 22th century digging into the old backup tapes of Reddit and attempt to rerun the system.
But ultimately, it's a problem that is needed to be addressed by the protocol and software with archive and preservation in mind.
---
BTW: A few weeks ago, I've written a lengthy comment on the fundamental conflict of history preservation and personal privacy, using the Usenet as an example, you may find it interesting.
Re: Internet Data Is Rotting
#100I’m okay with internet rot and you should be too. I’m not sure where we got the idea that “our data must be preserved forever”. This can be especially harmful for teens and young adults whose indiscretions now follow them forever. Think of the privilege you had when you were younger. You could do something stupid and nobody could whip out a high def camera to record it and make it part of your history forever. Let it…
I'm OK with it because otherwise you are whitewashing history. For example I have recordings of the Colbert report going back to ~2005. Some of his skits released during that time would be classified as "hate speech" in 2019. Of course he, and mainstream broadcasting companies would love it if you didn't think about that. There are plenty of news clips and interviews where mainstream politicians (on Left AND Right) c…