Live data from Hacker News

Internet Archive as a default host-of-record for startups

twitter.com

61–70 of 188 posts

Re: Internet Archive as a default host-of-record for startups

#62
post #9

The feature I most want from the Internet Archive is the ability to donate them an old domain name and enough cash to renew it for the next hundred years such that they can keep an archived version of a site available (without breaking any incoming links) for a very long time. They would also need to be able to handle legal administration costs of things like DMCA take-down notices, but I assume they already have to…

It costs the Internet Archive $2/GB to host content in perpetuity. They have a tool, Archive It, that will periodically crawl your site for archival purposes if you are not technical. For my needs, I run a report monthly for the content I’ve archived using my IA account to determine archived GBs, and then donate the amount needed to cover those costs. Consider reaching out to their patron services email address with…

> It costs the Internet Archive $2/GB to host content in perpetuity.

Do you have source/more info than that?

Lets say the internet archive is 100 PB [1], that's 100,000,000 GB [2], and at that rate it comes out to $200 million [3] for the whole thing forever. That's a lot of money, but also a lot less than I was expecting for something like that.

[1] https://www.protocol.com/internet-archive-preserving-future: "The web archive alone is about 45 petabytes — 4,500 terabytes — and the Internet Archive itself is about double that size (the group has other collections, like a huge database of educational films, music and even long-gone software programs)."

[2] https://www.google.com/search?q=100+petabytes+to+gb: "100 petabyte = 1e+8 gigabytes"

[3] https://www.google.com/search?q=1e%2B8+*+%242: "1e+8 * (US$ 2) = 200 million US$"

Re: Internet Archive as a default host-of-record for startups

#63
post #10

Earlier quoted context omitted.

Throughout human history, records have been forgotten, rewritten, changed, mutated, degraded, eroded away to nothingness. "The internet is forever" has always struck me as inhumane. Make a mistake or expose a weakness on the internet and it will always accompany you. It turns out that the internet is not always forever. I find that comforting.

Maybe humans should become more accommodating of past mistakes.

Suggest that to them. I'm sure they'll get right on it.

Re: Internet Archive as a default host-of-record for startups

#64
post #10
post #3

Personally, I think eternally archiving everything and infinitely available public data has been not-so-great. If this was an "archive with consent" sort of system, then sure. My response may be better summarized as, "Does IA support robots.txt, and if not why?"

Throughout human history, records have been forgotten, rewritten, changed, mutated, degraded, eroded away to nothingness. "The internet is forever" has always struck me as inhumane. Make a mistake or expose a weakness on the internet and it will always accompany you. It turns out that the internet is not always forever. I find that comforting.

I think there's inhumanity of a kind on both sides of this question. "Everything you've done will be forgotten, and no one will remember your name" is the sort of thing the bad guys say in movies. But that's what happens to most of us in the end. I think it's natural not to want that.

Re: Internet Archive as a default host-of-record for startups

#65
post #16

Earlier quoted context omitted.

You're consenting by posting it on public internet in the first place.

Posting something on the public internet is not consent for you to scrape it and post it on your own site forever. And requiring an explicit opt-in would basically mean no IA. To be clear, the IA is a positive, maybe even a great one. But it skirts by because most people don't care. (They did as you say post whatever on the public internet.) Add the facts that they're a non-profit, aren't trying to monetize their hos…

> Posting something on the public internet is not consent for you to scrape it and post it on your own site forever.

It effectively is. Your consent is not required, and people are doing far worse than just keeping it available (Clearview; there are also reports of people hoovering up encrypted data to crack in the coming decades when we're post-quantum).

This is no different than demanding people not keep track of anything else, and attacking archive.org might make you feel better, but that won't make anyone else stop.

Re: Internet Archive as a default host-of-record for startups

#66

Earlier quoted context omitted.

A far cheaper solution is a browser extension that looks up DNS differently based on the age of a the link. It wouldn't be hard to maintain a hand-crafted database of when domains are reused for something completely different, or even when the same conceptual website has breakages, and use that to choose between Internet Archive or live web accordingly. When one is browsing from an internet archive page, the date is…

Trying to figure out the age of links referenced against a hand crafted database (or trying to figure out if the current version is "too different" based on age automatically) using a browser extension only serves to create an unreliable solution for a few using the extension. Cheaper sure but it's also doing a whole lot less. Alternative content addressed systems may also work better for finding the content than som…

> using a browser extension only serves to create an unreliable solution for a few using the extension. Cheaper sure but it's also doing a whole lot less.

This is rather pessimistic thinking. The same money that goes into buying up domains could go into lobbying browsers to add this functionality be default.

> bulk of the problem space is in guaranteeing active hosting in a way viewable to viewers of the age will be available for many years not addressing the content.

You're moving the goal posts. I am not saying "IPFS means we don't need the internet archive". We absolute do need the internet archive. Content addressing helps by making archiving transparent, so the archival copy is not worse than the original.

Fundamentally, consumers producers or archivists may be the party most interested in the continued existence of some information at different moments in the lifespan of that information. Location-based addressing forces the producers to shoulder the burden of hosting, but content-based addressing allows the work to be distributed among those 3 however we see fit. Of course the burdened must still be barred! That doesn't mean the flexibility isn't extremely useful.

> On top of hosting and addressing the Internet Archive offers the ability to view old content on modern browsers even if modern browsers have 0 support for such content anymore (or if browsers ever had support at all even). Forward compatibility isn't something solved by a protocol.

Yeah that's great too, and again not something I am arguing is not good, or not necessary.

Re: Internet Archive as a default host-of-record for startups

#68

I think it's interesting to think about what we have lost because we couldn't keep everything from a 100 years ago and what society 100 years from now will be grateful we preserved. Off the top of my head, we lost a lot of common wisdom in dealing with the flu pandemic of 1918 because personal letters and most newspapers were not preserved. I think 100 years from now they might wish we had preserved more from margina…

Hard to know what will be of interest for future historians. Some things in which we place great value can be considered irrelevant, while some of our junk can become historical gold.

Re: Internet Archive as a default host-of-record for startups

#69
post #9

The feature I most want from the Internet Archive is the ability to donate them an old domain name and enough cash to renew it for the next hundred years such that they can keep an archived version of a site available (without breaking any incoming links) for a very long time. They would also need to be able to handle legal administration costs of things like DMCA take-down notices, but I assume they already have to…

A full-text search for the Wayback Machine would be my top feature request. It's not uncommon to lose the URL of a site and for active webpages to not have the URL of the old website. Plus I'm sure there are many interesting archived webpages I could find with a full-text search. I understand they've tried this or things like it a few times but they haven't ever kept the feature.

I imagine that the costs to run this would outweight the potential marketing benefits, but it'd be amazing to see Algolia take this project on to benefit everyone.

The Wayback Machine's data is ~20PB..? What is approximately the size of the indexable text (i.e. the text content of html pages, sans tags)? And what would the index size be like, approximately?

I imagine that creating (and maintaining, of course) the index would be the most time-consuming part? Is it at all possible to imagine hosting this index... somewhere... and doing sqlite http range-like queries on it..?

Would it be enough to have an index consist of a list of found words, and the related "document ids"? i.e. "apple" is in doc ids 1000, 2000, 3000, "banana" is in doc ids 2000, 4000, etc.?

And have separate docid -> archive.org url mapping?

Re: Internet Archive as a default host-of-record for startups

#70
post #59

I don't understand why Carmack thinks blockchain should be a component of this. Anyone care to elaborate on how that would make this easier/better?

I think he's referring to something like IPFS. https://en.wikipedia.org/wiki/InterPlanetary_File_System http://ipfs.io You can put the storage costs on the nodes because storage at archive.org's scale adds up, especially when it's run by volunteers.

It looks like you could use IPFS to accomplish this without using a blockchain.
Post reply on HN