Live data from Hacker News

Internet Data Is Rotting

theconversation.com

11–20 of 116 posts

Re: Internet Data Is Rotting

#11
post #2

> Then there is also a problem of software preservation: How can people today or in the future interpret those WordPerfect or WordStar files from the 1980s, when the original software companies have stopped supporting them or gone out of business? This issue in particular we have great solutions for (open formats / text), but they are of course less profitable than only-my-app-can-read-this formats.

FWIW those particular formats are widely understood even if they are proprietary (well, at least in WordStar's case). And as long as the software runs (be it natively or via an emulator or VM), you can always open and convert/print the files (e.g. you could use vDOS to run WordStar or whatever and use its printer emulator functionality with Windows' PDF printer to create a PDF from the WordStar files).

Re: Internet Data Is Rotting

#12

I read somewhere that the lifespan of the average hyperlink is only about two years. I count myself lucky I was introduced to the HTTRACK archiver program many years ago and thus have complete offline copies of many of my favorite websites of the early 00's.

Similarly, also even today i use wget --clone and the Firefox addon Save Page WE to save interesting pages (the latter works for single pages, but it is useful for blog articles, etc).

Re: Internet Data Is Rotting

#13
I’m okay with internet rot and you should be too. I’m not sure where we got the idea that “our data must be preserved forever”. This can be especially harmful for teens and young adults whose indiscretions now follow them forever.

Think of the privilege you had when you were younger. You could do something stupid and nobody could whip out a high def camera to record it and make it part of your history forever.

Let it rot.

Re: Internet Data Is Rotting

#14

I read somewhere that the lifespan of the average hyperlink is only about two years. I count myself lucky I was introduced to the HTTRACK archiver program many years ago and thus have complete offline copies of many of my favorite websites of the early 00's.

Can you give some examples of these 'favorite websites'? I'm interested in knowing what kind of website would be so interesting that I would want an entire offline copy of it. (Besides maybe Wikipedia)

Re: Internet Data Is Rotting

#15

I’m okay with internet rot and you should be too. I’m not sure where we got the idea that “our data must be preserved forever”. This can be especially harmful for teens and young adults whose indiscretions now follow them forever. Think of the privilege you had when you were younger. You could do something stupid and nobody could whip out a high def camera to record it and make it part of your history forever. Let it…

Letting MySpace rot is fine.

As the article points out, a more concerning issue is that “universities, governments and scientific societies are struggling to preserve scientific data”.

Re: Internet Data Is Rotting

#16
This needs to be solved on the protocol level. Of course, the players who have control over our protocols are exactly the people who don't want this to be solved at all.

The next best thing would be to redefine what "bookmarking" is. When I bookmark a page, I want it to be permanently stored on my local machine and full-text indexed. In fact, it's rather ridiculous that after 25 years browsers don't have anything of this sort. Unfortunately, the most popular browser in the world is controlled by the same people who control our protocols.

If I ever get the energy, I will attempt to write a browser extension for this.

Re: Internet Data Is Rotting

#17
post #16

This needs to be solved on the protocol level. Of course, the players who have control over our protocols are exactly the people who don't want this to be solved at all. The next best thing would be to redefine what "bookmarking" is. When I bookmark a page, I want it to be permanently stored on my local machine and full-text indexed. In fact, it's rather ridiculous that after 25 years browsers don't have anything of…

wget -r ?

Re: Internet Data Is Rotting

#18
post #17
post #16

This needs to be solved on the protocol level. Of course, the players who have control over our protocols are exactly the people who don't want this to be solved at all. The next best thing would be to redefine what "bookmarking" is. When I bookmark a page, I want it to be permanently stored on my local machine and full-text indexed. In fact, it's rather ridiculous that after 25 years browsers don't have anything of…

wget -r ?

No. There are some cases where it is useful to download many pages in a batch, but what I am talking about is, effectively, partial local replication. Bookmarking I describe should create a tiny (but useful) version of the web on your computer.

It needs to be seamless. It needs to be searchable. It would be incredibly useful if it would capture relationships between pages (links) in addition to pages themselves to navigate offline.

I can see many use cases for a kind of "super bookmark mode" where the browser automatically stores all the pages you visit within a certain domain.

Re: Internet Data Is Rotting

#19
post #18
post #17

Earlier quoted context omitted.

wget -r ?

No. There are some cases where it is useful to download many pages in a batch, but what I am talking about is, effectively, partial local replication. Bookmarking I describe should create a tiny (but useful) version of the web on your computer. It needs to be seamless. It needs to be searchable. It would be incredibly useful if it would capture relationships between pages (links) in addition to pages themselves to na…

[deleted]

Re: Internet Data Is Rotting

#20
We should aim at browsing the Internet by date. We're moving everything there ignoring the fact that as it is there is no built-in permanence. We're accustomed to that after Gutenberg: it wasn't that easy losing every copy of an important document. Now it is, things disappear and we're getting into a drifting cultural bubble impossible to trace back.

The Internet Archive is doing God's work, but it's not enough. If you don't have the URL of a site that is gone, you probably won't find any reference to it after every online hyperlink to it has disappeared as well. It might become then inaccessible after a while, stored yet gone anyway.

Post reply on HN