Live data from Hacker News

Internet Archive as a default host-of-record for startups

twitter.com

31–40 of 188 posts

Re: Internet Archive as a default host-of-record for startups

#31
post #24

Earlier quoted context omitted.

Public data is... public. No one should be stopped from saving a public page and nothing should stop Internet Archive, be it robot or human. There is if course a need for removal of archived content infringing on someone's rights, whatever that might be, but "archive with consent" will fail for the goal of preserving culture. I think it's worrying that some online newspapers enacted archive blockers or IA needing DMC…

Content posted on a web site is NOT “public” (domain), it is (in the US) automatically copyrighted to the author, unless they specifically waive those rights. Just because you can see it through a browser doesn’t in any way mean you can make it yours and do what you want with it.

We store books in public libraries even if they aren't public domain and authors can do nothing to prevent them from doing so.

Re: Internet Archive as a default host-of-record for startups

#32
post #3

Personally, I think eternally archiving everything and infinitely available public data has been not-so-great. If this was an "archive with consent" sort of system, then sure. My response may be better summarized as, "Does IA support robots.txt, and if not why?"

There are clearly some things we, and the author, wants to persist, and yet the internet fails to do so. Trying to muddle in privacy concerns with the accumulation of public knowledge undermines the whole concept of a shared society, or there being even the potential of accumulating "progress" in the first place. Also, historically speaking, people with means used to save their letters for posterity, which proved to…

> Also, historically speaking, people with means used to save their letters for posterity, which proved to be a very valuable resource for future academics

Your example is an example of choice or consent. They also had the option to burn their hand written books and scrolls down periodically. Systems like IA take that choice away.

Re: Internet Archive as a default host-of-record for startups

#33
post #10
post #3

Personally, I think eternally archiving everything and infinitely available public data has been not-so-great. If this was an "archive with consent" sort of system, then sure. My response may be better summarized as, "Does IA support robots.txt, and if not why?"

Throughout human history, records have been forgotten, rewritten, changed, mutated, degraded, eroded away to nothingness. "The internet is forever" has always struck me as inhumane. Make a mistake or expose a weakness on the internet and it will always accompany you. It turns out that the internet is not always forever. I find that comforting.

Maybe humans should become more accommodating of past mistakes.

Re: Internet Archive as a default host-of-record for startups

#34
post #27
post #9

The feature I most want from the Internet Archive is the ability to donate them an old domain name and enough cash to renew it for the next hundred years such that they can keep an archived version of a site available (without breaking any incoming links) for a very long time. They would also need to be able to handle legal administration costs of things like DMCA take-down notices, but I assume they already have to…

I agree, i'm a passionate photographer and i could pay good money to know that my pictures could be seen long time after my death. Maybe startups exists that do this, but they will die, i need something with enough critical mass that i can trust.

Perhaps have a look at Arweave (https://www.arweave.org).

Re: Internet Archive as a default host-of-record for startups

#35
post #24

Earlier quoted context omitted.

Public data is... public. No one should be stopped from saving a public page and nothing should stop Internet Archive, be it robot or human. There is if course a need for removal of archived content infringing on someone's rights, whatever that might be, but "archive with consent" will fail for the goal of preserving culture. I think it's worrying that some online newspapers enacted archive blockers or IA needing DMC…

Content posted on a web site is NOT “public” (domain), it is (in the US) automatically copyrighted to the author, unless they specifically waive those rights. Just because you can see it through a browser doesn’t in any way mean you can make it yours and do what you want with it.

Absolutely true. "unless they specifically waive those rights" - if archiving entailed contacting the owner with a legal archive request, we would have archived basically nothing. Luckily there are exceptions for Internet Archive in place. My point is, if "by consent" was the requirement to archive information, we would have archived nothing.

Re: Internet Archive as a default host-of-record for startups

#36
post #24

Earlier quoted context omitted.

Content posted on a web site is NOT “public” (domain), it is (in the US) automatically copyrighted to the author, unless they specifically waive those rights. Just because you can see it through a browser doesn’t in any way mean you can make it yours and do what you want with it.

We store books in public libraries even if they aren't public domain and authors can do nothing to prevent them from doing so.

If you buy the book, you have an authorization to do so. But you can't put copies of it.

Re: Internet Archive as a default host-of-record for startups

#37
post #9

The feature I most want from the Internet Archive is the ability to donate them an old domain name and enough cash to renew it for the next hundred years such that they can keep an archived version of a site available (without breaking any incoming links) for a very long time. They would also need to be able to handle legal administration costs of things like DMCA take-down notices, but I assume they already have to…

A far cheaper solution is a browser extension that looks up DNS differently based on the age of a the link. It wouldn't be hard to maintain a hand-crafted database of when domains are reused for something completely different, or even when the same conceptual website has breakages, and use that to choose between Internet Archive or live web accordingly. When one is browsing from an internet archive page, the date is…

I made something very similar, though instead of age, it enforces page authorship for links (using PGP signatures). https://webverify.jahed.dev/

Re: Internet Archive as a default host-of-record for startups

#38
post #2

Correct me if I'm wrong, but isn't this problem the ideal use case for projects like IPFS? Anyone interested to preserve the content can join as a node to balance the load, right? And if so, why don't we see widespread adoption?

IPFS only does addressability, it doesn't provide storage. You could use a decentralized storage network like Arweave, Filecoin, or Sia.

https://www.arweave.org/

https://www.filecoin.com/

http://sia.tech/

Re: Internet Archive as a default host-of-record for startups

#39
post #9

The feature I most want from the Internet Archive is the ability to donate them an old domain name and enough cash to renew it for the next hundred years such that they can keep an archived version of a site available (without breaking any incoming links) for a very long time. They would also need to be able to handle legal administration costs of things like DMCA take-down notices, but I assume they already have to…

It costs the Internet Archive $2/GB to host content in perpetuity. They have a tool, Archive It, that will periodically crawl your site for archival purposes if you are not technical.

For my needs, I run a report monthly for the content I’ve archived using my IA account to determine archived GBs, and then donate the amount needed to cover those costs.

Consider reaching out to their patron services email address with any questions.

Edit: $2/GB citation: https://help.archive.org/hc/en-us/articles/360014755952-Arch...

Re: Internet Archive as a default host-of-record for startups

#40

I think it's interesting to think about what we have lost because we couldn't keep everything from a 100 years ago and what society 100 years from now will be grateful we preserved. Off the top of my head, we lost a lot of common wisdom in dealing with the flu pandemic of 1918 because personal letters and most newspapers were not preserved. I think 100 years from now they might wish we had preserved more from margina…

An excellent place to start is talking to your parents and grandparents and recording their history and stories online.
Post reply on HN