Live data from Hacker News

Screenshots Forever and Ever Until You Can’t Stand it

ascii.textfiles.com

31–40 of 41 posts

Re: Screenshots Forever and Ever Until You Can’t Stand it

#31
post #28

Earlier quoted context omitted.

I have a few one-off remaining records and I'm very conscious of the responsibility that goes with that. I've never played them and the only time they will be played is when there is a very high grade digitization done at the same time. I've been looking at ways to digitize them without having to actually mechanically play them (using lasers) but those systems are expensive! ( http://www.elpj.com/product/price.php )

Well, sure, buying one is expensive, but you're just looking to use one. You only need one to share amongst whomever else you can find who has this need, and the Internet is really good at bringing those sorts of people together. You could Kickstart your way up to owning one of these, just as one idea. (Or work with somebody else to take it on, etc. etc., whatever.)

Kickstarter is an interesting option, never thought of that.

You could jumpstart a service that digitizes rare recordings as a service and make them available. Likely there'd be all sorts of copyright issues with a service like that but the basic principle is useful.

Re: Screenshots Forever and Ever Until You Can’t Stand it

#32

Earlier quoted context omitted.

Do you mean to tell me that those crates of my dad's 78 rpm records sitting in my basement might actually be worth something? I was thinking about playing them on a 33.3 rpm turntable into my PC, then using Audacity to speed them up to 78, just to be able to hear the music again, and by association, remember my dad. Geeze, there must be at least a hundred of 'em down there, at least!

Google the work of philip jeck (samples on youtube) for hugely emotional use of the old sounds. Try and get hold of an actual 78 rpm turntable, the old dansettes could do those speeds as could the cheaper separate turntables that had ceramic cartridges with two 'sides' (78/33). Enjoy

Rega make a dedicated 78rpm turntable:

http://www.rega.co.uk/rp78.html

Re: Screenshots Forever and Ever Until You Can’t Stand it

#34
post #28

Earlier quoted context omitted.

Well, sure, buying one is expensive, but you're just looking to use one. You only need one to share amongst whomever else you can find who has this need, and the Internet is really good at bringing those sorts of people together. You could Kickstart your way up to owning one of these, just as one idea. (Or work with somebody else to take it on, etc. etc., whatever.)

Kickstarter is an interesting option, never thought of that. You could jumpstart a service that digitizes rare recordings as a service and make them available. Likely there'd be all sorts of copyright issues with a service like that but the basic principle is useful.

Given the sort of thing we're talking about, I think a "ask for forgiveness rather than permission" approach is at least morally justifiable.

I have to admit that before I'd start doing this at scale, I would want some sort of LLC-esque shield in place.

Re: Screenshots Forever and Ever Until You Can’t Stand it

#35
post #23

Earlier quoted context omitted.

Imo, a better approach than nuking old data, could be to keep the data, but not show it.

I may be mistaken, but I think I read somewhere that's exactly what they do. But don't take my word for it.

You are not mistaken. The Internet Archive does not delete or "nuke" the data that is blocked by a robots.txt. Even though cough some people believe so (see parent thread).

Source: IA staffers.

Re: Screenshots Forever and Ever Until You Can’t Stand it

#36
post #9

Earlier quoted context omitted.

A few years ago I talked to an IA engineer, who said they were planning on dealing with this by not crawling sites whose nameservers were known to point to a domain parking company. The idea was that if they never retrieved the robots.txt, they wouldn't retroactively apply it. I don't know if that filtering out of parking nameservers ever happened, and it wouldn't help for parked domains whose robots.txt they'd alrea…

But why retroactively remove the data? The original owner was fine with holding it, why should the snapshot be deleted because a completely different person wants his completely different website to not be crawled?

It's hard for a bot to understand concepts of 'owner' and 'completely different person' based on the data they have available. Companies can use this robots.txt feature to un-index old marketing content after a re-branding, for example. Or after an acquisition.

Re: Screenshots Forever and Ever Until You Can’t Stand it

#37
post #36

Earlier quoted context omitted.

But why retroactively remove the data? The original owner was fine with holding it, why should the snapshot be deleted because a completely different person wants his completely different website to not be crawled?

It's hard for a bot to understand concepts of 'owner' and 'completely different person' based on the data they have available. Companies can use this robots.txt feature to un-index old marketing content after a re-branding, for example. Or after an acquisition.

Sure, but, surely, the bot has timestamps saying "robots.txt allowed me to keep these documents last time I spidered them". Why do they have to be retroactively removed? robots.txt only disallows spidering, it doesn't mandate that you should delete all the data you've already spidered.

Re: Screenshots Forever and Ever Until You Can’t Stand it

#38
post #3

I've been thinking a lot the past few weeks about the preservation of art and goods in general. I think it comes from two main sources: 1. Maciej Ceglowski (`idlewords) has been talking a lot recently about link rot. Recently, he found that 25% of items pinned a mere five years ago were dead links and 17% from only three years ago were dead [^1]. That's an incredibly high rate -- and, selfishly, one I'm noticing as m…

Do you mean to tell me that those crates of my dad's 78 rpm records sitting in my basement might actually be worth something? I was thinking about playing them on a 33.3 rpm turntable into my PC, then using Audacity to speed them up to 78, just to be able to hear the music again, and by association, remember my dad. Geeze, there must be at least a hundred of 'em down there, at least!

Please find proper equipment to play them.

Using a microgroove stylus designed for vinyl will not do your shellac records any favours. Also the recordings will sound terrible.

There are several modern cartridges that support 78rpm stylii, and lots of vintage kit available cheap.

Also careful, they shatter somewhat easily, and like to get mouldy.

Re: Screenshots Forever and Ever Until You Can’t Stand it

#39
post #16
post #5

Earlier quoted context omitted.

> The nice things about IA links is you can pretty reasonably assume that they won't suffer from link rot, right? The catch there being, the Internet Archive retroactively respects robots.txt that forbid crawling, so if someone gets control of a domain they can block the archived pages. This is a big problem with lapsed domains that get swept under the umbrella of a holding company that has lots of domains pointing t…

This is a huge problem. Sites like NASA's NTRS are retroactively blocked.[1] It's not clear which user agents one must allow in robots.txt. NTRS allows archive.org_bot, but apparently ia_archiver is also needed. At some point the allow directive in the NTRS robots.txt[2] no longer matched, nuking all historical data. 1. See http://web.archive.org/web/20121029225832/http://ntrs.nasa.g... for an example. 2. http://ntrs…

It's not been nuked, merely hidden.

Re: Screenshots Forever and Ever Until You Can’t Stand it

#40
post #36

Earlier quoted context omitted.

It's hard for a bot to understand concepts of 'owner' and 'completely different person' based on the data they have available. Companies can use this robots.txt feature to un-index old marketing content after a re-branding, for example. Or after an acquisition.

Sure, but, surely, the bot has timestamps saying "robots.txt allowed me to keep these documents last time I spidered them". Why do they have to be retroactively removed? robots.txt only disallows spidering, it doesn't mandate that you should delete all the data you've already spidered.

Because most of the problems come from people who want to hide old material that they didn't realize was being indexed. The automatic behavior is simple and easy to implement, and doesn't require any human intervention.
Post reply on HN