Live data from Hacker News

Screenshots Forever and Ever Until You Can’t Stand it

ascii.textfiles.com

1–10 of 41 posts

Re: Screenshots Forever and Ever Until You Can’t Stand it

#2
Ah ZX Spectrum that takes me back. I learned assembler on it, BASIC, had a Pascal and C compiler even (the last two had to load from tape). Made one mistake and had to reload the whole thing again and wait 5 minutes or so.

It was amazing how there was this closeness to the machine, you boot right into the programming environment and had to type a command to load a game or do anything.

Some of the games I remember were exceptional, Elite was one of them. Just thinking about the ability to pack everything in 48K of memory.

Re: Screenshots Forever and Ever Until You Can’t Stand it

#3
I've been thinking a lot the past few weeks about the preservation of art and goods in general. I think it comes from two main sources:

1. Maciej Ceglowski (`idlewords) has been talking a lot recently about link rot. Recently, he found that 25% of items pinned a mere five years ago were dead links and 17% from only three years ago were dead [^1]. That's an incredibly high rate -- and, selfishly, one I'm noticing as my bookmark folders for recipes (RIP BroEats.com) and designs and other things is filled with more and more duds.

2. I've been reading Do Not Sell At Any Price [^2], a book about the subculture of 78rpms. These are records that are so rare and so -- for lack of a better term -- unwanted by the vast majority of the music-listening populace that the act of collecting them is less about hoarding and more about preservation. To quote one of the characters in the book (roughly from memory):

"It's a weird feeling, holding this thing in your hand and knowing that you could break the song," he said. "I snap this record in half and this song is lost forever. It's a lot of responsibility, and sometimes I think that's why I take it so seriously."

I can honestly say that, prior to reading this article, I had no idea what a ZX Spectrum was. Now, after some digging, I do -- and I still have no desire to play one, obtain one, or hold onto it in any meaningful way. (And seeing as I'm usually on the weirdly attached end of the spectrum with these kinds of things, I doubt I'm the only one.) But I'm struck by how important it is to hold onto these things, even if its in a cardboard box in a forgotten closet somewhere or a link on the Internet Archive that gets clicked once every couple decades.

I'm not positing that there will ever be a point in time that someone has the hankering to play ZX Spectrum Xtreme Chess, but I think there's inherent value in preserving this ecosystem -- something of a testament to the people who made it, the people who played it, the novelty that at one point in time there were five million living rooms with this machine in it.

The Web turned 25 this year, and it's already coming down with acute cases of memory loss. I'm hoping that by the time it hits fifty, the problem won't have gotten worse -- it will have gotten much, much better, not just with URLs but with remembering the time when people played 3D StarFighter by the Oliver Twins. [^4]

(This is a very roundabout way of saying the following: Jason, you are completely awesome for doing this, and thanks for sharing it with us.)

[^1] https://twitter.com/Pinboard/status/501406303747457024

[^2] http://www.amazon.com/Not-Sell-Any-Price-Obsessive-ebook/dp/...

[^3] The nice things about IA links is you can pretty reasonably assume that they won't suffer from link rot, right?

[^4] https://www.flickr.com/photos/textfiles/14771850314/in/set-7...

Re: Screenshots Forever and Ever Until You Can’t Stand it

#4
One of my first websites was an emulation site in 1999/2000. SNES was by far the most popular category on the site so I had the bright idea to play 2-3 minutes of every single game in order to take a screenshot of the game-play. Three months later I was finally done. Then I lost everything in hard drive crash before I could push the new version of my site live. Dammit.

Re: Screenshots Forever and Ever Until You Can’t Stand it

#5
post #3

I've been thinking a lot the past few weeks about the preservation of art and goods in general. I think it comes from two main sources: 1. Maciej Ceglowski (`idlewords) has been talking a lot recently about link rot. Recently, he found that 25% of items pinned a mere five years ago were dead links and 17% from only three years ago were dead [^1]. That's an incredibly high rate -- and, selfishly, one I'm noticing as m…

> The nice things about IA links is you can pretty reasonably assume that they won't suffer from link rot, right?

The catch there being, the Internet Archive retroactively respects robots.txt that forbid crawling, so if someone gets control of a domain they can block the archived pages. This is a big problem with lapsed domains that get swept under the umbrella of a holding company that has lots of domains pointing to the same content, with a blanket robots.txt.

Re: Screenshots Forever and Ever Until You Can’t Stand it

#6
post #5
post #3

I've been thinking a lot the past few weeks about the preservation of art and goods in general. I think it comes from two main sources: 1. Maciej Ceglowski (`idlewords) has been talking a lot recently about link rot. Recently, he found that 25% of items pinned a mere five years ago were dead links and 17% from only three years ago were dead [^1]. That's an incredibly high rate -- and, selfishly, one I'm noticing as m…

> The nice things about IA links is you can pretty reasonably assume that they won't suffer from link rot, right? The catch there being, the Internet Archive retroactively respects robots.txt that forbid crawling, so if someone gets control of a domain they can block the archived pages. This is a big problem with lapsed domains that get swept under the umbrella of a holding company that has lots of domains pointing t…

Have you heard of other caches/archives (e.g. Google) applying the same retroactive policy? Presmably IA has no way of finding out that domain ownership has changed. I wonder if they are applying this policy to pages referenced by Wikipedia, http://blog.archive.org/2013/10/25/fixing-broken-links/

The safest archive of a web page is a local PDF.

Re: Screenshots Forever and Ever Until You Can’t Stand it

#7
post #3

I've been thinking a lot the past few weeks about the preservation of art and goods in general. I think it comes from two main sources: 1. Maciej Ceglowski (`idlewords) has been talking a lot recently about link rot. Recently, he found that 25% of items pinned a mere five years ago were dead links and 17% from only three years ago were dead [^1]. That's an incredibly high rate -- and, selfishly, one I'm noticing as m…

A good source of archival resources is http://fileformat.info

Re: Screenshots Forever and Ever Until You Can’t Stand it

#8
post #5
post #3

I've been thinking a lot the past few weeks about the preservation of art and goods in general. I think it comes from two main sources: 1. Maciej Ceglowski (`idlewords) has been talking a lot recently about link rot. Recently, he found that 25% of items pinned a mere five years ago were dead links and 17% from only three years ago were dead [^1]. That's an incredibly high rate -- and, selfishly, one I'm noticing as m…

> The nice things about IA links is you can pretty reasonably assume that they won't suffer from link rot, right? The catch there being, the Internet Archive retroactively respects robots.txt that forbid crawling, so if someone gets control of a domain they can block the archived pages. This is a big problem with lapsed domains that get swept under the umbrella of a holding company that has lots of domains pointing t…

What you say is true for sites mirrored by IA's Wayback Machine, though my understanding is they retain the data in case the robots.txt is lifted later on.

Linking to media uploaded to the main archive itself should be safer, though.

Re: Screenshots Forever and Ever Until You Can’t Stand it

#9
post #5
post #3

I've been thinking a lot the past few weeks about the preservation of art and goods in general. I think it comes from two main sources: 1. Maciej Ceglowski (`idlewords) has been talking a lot recently about link rot. Recently, he found that 25% of items pinned a mere five years ago were dead links and 17% from only three years ago were dead [^1]. That's an incredibly high rate -- and, selfishly, one I'm noticing as m…

> The nice things about IA links is you can pretty reasonably assume that they won't suffer from link rot, right? The catch there being, the Internet Archive retroactively respects robots.txt that forbid crawling, so if someone gets control of a domain they can block the archived pages. This is a big problem with lapsed domains that get swept under the umbrella of a holding company that has lots of domains pointing t…

A few years ago I talked to an IA engineer, who said they were planning on dealing with this by not crawling sites whose nameservers were known to point to a domain parking company. The idea was that if they never retrieved the robots.txt, they wouldn't retroactively apply it. I don't know if that filtering out of parking nameservers ever happened, and it wouldn't help for parked domains whose robots.txt they'd already retrieved, but but it would help with domains that lapse in the future.

Re: Screenshots Forever and Ever Until You Can’t Stand it

#10
post #5

Earlier quoted context omitted.

> The nice things about IA links is you can pretty reasonably assume that they won't suffer from link rot, right? The catch there being, the Internet Archive retroactively respects robots.txt that forbid crawling, so if someone gets control of a domain they can block the archived pages. This is a big problem with lapsed domains that get swept under the umbrella of a holding company that has lots of domains pointing t…

Have you heard of other caches/archives (e.g. Google) applying the same retroactive policy? Presmably IA has no way of finding out that domain ownership has changed. I wonder if they are applying this policy to pages referenced by Wikipedia, http://blog.archive.org/2013/10/25/fixing-broken-links/ The safest archive of a web page is a local PDF.

> The safest archive of a web page is a local PDF.

Minor quibble: The safest archive of a web page is a local WARC archive:

http://www.archiveteam.org/index.php?title=Wget_with_WARC_ou...

Post reply on HN