Live data from Hacker News

Internet Archive as a default host-of-record for startups

twitter.com

21–30 of 188 posts

Re: Internet Archive as a default host-of-record for startups

#21
post #3

Personally, I think eternally archiving everything and infinitely available public data has been not-so-great. If this was an "archive with consent" sort of system, then sure. My response may be better summarized as, "Does IA support robots.txt, and if not why?"

You're consenting by posting it on public internet in the first place.

Do you think most people who post publicly on the internet would agree if asked? I think most people would like to have a choice to make old stuff disappear. If most people think so, THAT should be the rule. You might not like that and argue that there is no way to enforce it, but that does not mean it is a good rule to assume consent.

Re: Internet Archive as a default host-of-record for startups

#22
post #9

The feature I most want from the Internet Archive is the ability to donate them an old domain name and enough cash to renew it for the next hundred years such that they can keep an archived version of a site available (without breaking any incoming links) for a very long time. They would also need to be able to handle legal administration costs of things like DMCA take-down notices, but I assume they already have to…

A full-text search for the Wayback Machine would be my top feature request. It's not uncommon to lose the URL of a site and for active webpages to not have the URL of the old website. Plus I'm sure there are many interesting archived webpages I could find with a full-text search.

I understand they've tried this or things like it a few times but they haven't ever kept the feature.

Re: Internet Archive as a default host-of-record for startups

#24
post #3

Personally, I think eternally archiving everything and infinitely available public data has been not-so-great. If this was an "archive with consent" sort of system, then sure. My response may be better summarized as, "Does IA support robots.txt, and if not why?"

Public data is... public. No one should be stopped from saving a public page and nothing should stop Internet Archive, be it robot or human. There is if course a need for removal of archived content infringing on someone's rights, whatever that might be, but "archive with consent" will fail for the goal of preserving culture. I think it's worrying that some online newspapers enacted archive blockers or IA needing DMC…

Content posted on a web site is NOT “public” (domain), it is (in the US) automatically copyrighted to the author, unless they specifically waive those rights. Just because you can see it through a browser doesn’t in any way mean you can make it yours and do what you want with it.

Re: Internet Archive as a default host-of-record for startups

#25
Imagine someone building this for SaaS hosting--a perma-Heroku, or something like it. That's actually a huge value-add. Suddenly tinyStartupA doesn't need to convince largeCorpB that it's going to be around for forever. The service can exist in perpetuity without the company.

Complex repercussions obviously around acquisition, IP, and other business dimensions however. Maybe unworkable even. But I think there's a world where this actually exists and lowers the barrier to building business-critical software and selling to companies that need a 50-year commitment to risk you.

Re: Internet Archive as a default host-of-record for startups

#26
post #17

Earlier quoted context omitted.

I was involved with planning https://nlnet.nl/project/SoftwareHeritage-P2P/ for just this reason --- hopefully we will finally be able to start work on it sometime too far off. Indeed the real challenge of archival is not loosing the stuff, by making sure that people can still find the stuff. "Orphaned" information that no one knows exists, or is bothering to interact with, isn't that valuable compared to resources t…

The great thing about location-based addressing is that an archive of the set of known locations is not subject to the same ownership rules as the canonical live version of those addresses. A document listing all Geocities URLs can be placed in content-addressed storage without needing geocities.com to be owned by the party that emplaces that document. And a chain can be maintained such that people are incentivized t…

> The great thing about location-based addressing is that an archive of the set of known locations is not subject to the same ownership rules as the canonical live version of those addresses.

Erm, to me this sounds like putting up with link rot as hack around bad IP law? There are already IP exceptions for preservation. And if content-addressing was the norm, geocities-type sites might bow to market pressure to not "own" the content, but merely have some some sort of license for being the exclusive pinning service and running the ads or whatever. This is like avoiding the problem where your the rent on your current apartment doesn't fall as much as the market writ large because your landlord knows moving is not free.

Re: Internet Archive as a default host-of-record for startups

#27
post #9

The feature I most want from the Internet Archive is the ability to donate them an old domain name and enough cash to renew it for the next hundred years such that they can keep an archived version of a site available (without breaking any incoming links) for a very long time. They would also need to be able to handle legal administration costs of things like DMCA take-down notices, but I assume they already have to…

I agree, i'm a passionate photographer and i could pay good money to know that my pictures could be seen long time after my death. Maybe startups exists that do this, but they will die, i need something with enough critical mass that i can trust.

Re: Internet Archive as a default host-of-record for startups

#28
post #2

Correct me if I'm wrong, but isn't this problem the ideal use case for projects like IPFS? Anyone interested to preserve the content can join as a node to balance the load, right? And if so, why don't we see widespread adoption?

It is -- and Brewster Kahle and the Archive have been thinking about this for a long while (see this talk from him five years ago: https://archive.org/details/LockingTheWebOpen_2016 ). The model you can think for this would be to have as the Archive as the "node of last resort" of content-addressable storage, making sure there's always one node up with the content you want.

The incentive challenges are making sure that the average number of nodes is more than one, because, as Brewster likes to say, "libraries burn; it's what they do", plus all the traditional challenges of maintaining a commons at high levels of resilience. Once you have data on a network like IPFS, we can use a number of incentive models to make sure it stays there, including charitable projects like the Archive, government support (archives are traditionally state projects -- if every country's archive was pinning this content, it would be far more resilient), and decentralized incentive frameworks like Filecoin.

(Disclosure: I work for the Filecoin Foundation; in our decentralized preservation work, we've funded the Internet Archive's work in this area, though I should emphasise that IA works with a lot of different decentralizing technologies through their https://getdweb.net/ community.)

Re: Internet Archive as a default host-of-record for startups

#29
post #15
post #2

Correct me if I'm wrong, but isn't this problem the ideal use case for projects like IPFS? Anyone interested to preserve the content can join as a node to balance the load, right? And if so, why don't we see widespread adoption?

IPFS does the opposite, right? It doesn't guarantee the archive is available, which is what Carmack is asking for. Incentives to scale bandwidth with need already exist as long as you have he data at all. That is to say, IPFS doesn't help if the desire blooms after the nodes dry up. Things could still be lost.

The point of IPFS is not keep the data archived, but to allows users to no care who does the archiving.

Concretely, this would be to skip the "many people on encountering a dead URL don't bother to try the internet archive" problem.

Re: Internet Archive as a default host-of-record for startups

#30
I think it's interesting to think about what we have lost because we couldn't keep everything from a 100 years ago and what society 100 years from now will be grateful we preserved.

Off the top of my head, we lost a lot of common wisdom in dealing with the flu pandemic of 1918 because personal letters and most newspapers were not preserved. I think 100 years from now they might wish we had preserved more from marginal and/or world communities. What folk wisdom is being lost? Perhaps we need to expand our definition of what is worth saving.

Post reply on HN