Live data from Hacker News

Internet Archive as a default host-of-record for startups

twitter.com

111–120 of 188 posts

Re: Internet Archive as a default host-of-record for startups

#111
post #32

Earlier quoted context omitted.

> Also, historically speaking, people with means used to save their letters for posterity, which proved to be a very valuable resource for future academics Your example is an example of choice or consent. They also had the option to burn their hand written books and scrolls down periodically. Systems like IA take that choice away.

There is no reason to do person communication with staticish websites. Don't foist the problems of social media onto the Internet Archive.

I'm not sure I understand. robots.txt was a matter of consent and was standardized well over twenty years ago. The IA willfully ignores it because they believe it interferes with their mission. This was long before social media, when static websites were more dominant than dynamic ones.

Source: https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea...

Re: Internet Archive as a default host-of-record for startups

#112
post #42
post #21

Earlier quoted context omitted.

Do you think most people who post publicly on the internet would agree if asked? I think most people would like to have a choice to make old stuff disappear. If most people think so, THAT should be the rule. You might not like that and argue that there is no way to enforce it, but that does not mean it is a good rule to assume consent.

There are two things: 1.) Most people won't opt-in because a significant majority accept defaults and don't opt into most things. 2.) For people like yourself probing a bit deeper, you might well ask whether you really want to give up your ability to decide you don't want something you thought was so funny when you wrote it at 20 now that you're a politician running for office or up for a political appointment.

I mean even honoring the opt-out of robots.txt would be fantastic. As another commenter pointed out they willfully ignore it: https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea...

It's fairly unethical.

Re: Internet Archive as a default host-of-record for startups

#113
post #86

Earlier quoted context omitted.

> The feature I most want from the Internet Archive The feature I most want from IA is a streamlined system to delete content they have archived on domains that I own, including a proper privacy law compliance effort on their part. They have intentionally made it a difficult, manual process to get content removed. They operate as a de facto malicious crawler. They massively violate GDPR with how they operate and few…

Shouldn't it be hard to delete things from a library and historical-archive? If you had to choose between the GDPR, & an accurate historical record, which would you prefer?

We won't. Pretty much the only way things leave archives of record (which is what IA is trying to be) are through Acts of God (if the archive burns down/all of IA's servers are taken out in a mass alien EMP attack).

An author suggesting that the LoC remove their copy of a book/other work (including digital works) because they want to unpublish it would not fly.

The parent comment has an issue with anything on the Web not hidden in some way being considered 'public' and 'published', but that would be something that would require international cooperation to hash out.

Re: Internet Archive as a default host-of-record for startups

#114

I don't understand why Carmack thinks blockchain should be a component of this. Anyone care to elaborate on how that would make this easier/better?

.... He says it right after "to make internet applications that could outlive companies". If it's on the blockchain it doesn't matter if the company storing all of the archives shuts down, the content would still exist, forever, until their is a network running the chain. I suppose something like Torrent could be used?

He has a point, but it's that private companies shouldn't archives of record.

I actually think using blockchain for things like ensuring providence is interesting, since in archives being able to have a clean record of what happened to a piece is VERY useful. It just won't earn a ton of money, so we'll need to wait for the capitalism to burn off to see more not-for-profit uses.

Re: Internet Archive as a default host-of-record for startups

#115
post #24

Earlier quoted context omitted.

Public data is... public. No one should be stopped from saving a public page and nothing should stop Internet Archive, be it robot or human. There is if course a need for removal of archived content infringing on someone's rights, whatever that might be, but "archive with consent" will fail for the goal of preserving culture. I think it's worrying that some online newspapers enacted archive blockers or IA needing DMC…

Content posted on a web site is NOT “public” (domain), it is (in the US) automatically copyrighted to the author, unless they specifically waive those rights. Just because you can see it through a browser doesn’t in any way mean you can make it yours and do what you want with it.

It doesn't matter.

Archives are exempt from being forbidden to create copies due to copyright infringement. The Library of Congress can make all the copies it wants, it just can't SELL them.

Now, there is a question whether a private company should legally be able to BE an archive of record, but as of now there's no legal reason they can't be, I believe. So it's legal.

Re: Internet Archive as a default host-of-record for startups

#116
post #36

Earlier quoted context omitted.

We store books in public libraries even if they aren't public domain and authors can do nothing to prevent them from doing so.

If you buy the book, you have an authorization to do so. But you can't put copies of it.

You do if you're an archive, actually.

Re: Internet Archive as a default host-of-record for startups

#117
post #16

Earlier quoted context omitted.

Posting something on the public internet is not consent for you to scrape it and post it on your own site forever. And requiring an explicit opt-in would basically mean no IA. To be clear, the IA is a positive, maybe even a great one. But it skirts by because most people don't care. (They did as you say post whatever on the public internet.) Add the facts that they're a non-profit, aren't trying to monetize their hos…

It can be. The reason libraries and other archives have special rights is because they fought for them against the express wishes of people who sold paper. There are no arguments made against archive.org that weren't also made against libraries.

Thank you.

These rights are also under constant attack: It's normal to charge libraries exorbitant prices for digital materials compared to their analogue counterparts, for example.

Re: Internet Archive as a default host-of-record for startups

#118
post #74

I don't understand why Carmack thinks blockchain should be a component of this. Anyone care to elaborate on how that would make this easier/better?

"Immutable and existing in perpetuity" are good qualities for an archive service, and that's at least the idea with a blockchain.

It is interesting though, what happens if you put so much data into a blockchain. I guess only a couple of nodes would want to verify the validity of the chain (because you need all the data to do it). And would those nodes really be more likely to keep the data then the situation we are in now?

I guess after some time the nodes would agree on the hash and throw away the data because it would cost too much to store.

Re: Internet Archive as a default host-of-record for startups

#119
post #87

The idea is more interesting when you think about scrapers. Take any ecommerce website, there are several scrapers that download all pages every hours, it would be more efficient if a provider had a live copy of the website and then serve the requests to the scrapers or could even send webhook. A website could handle tons of scrappers without having high bandwidth, only the provider will need high bandwidth. The issu…

(in the nicest way and since your post is recent, please edit s/scrappers/scrapers! It very much changed how I read it first time around -- I thought you were referring to a type of failed startup!)

Thanks, I need to improve my english haha

Re: Internet Archive as a default host-of-record for startups

#120
What compression does IA use to store websites? Using a 2x better compression will allow them to store 2x more websites/content.

I am doing some compression research, and would love to help IA in any way I can. There are some amazing SOTA compression algorithms available now.

And if IA ignores images/video, and focuses only on text, they can store an insane amount of websites at a very low cost.

Post reply on HN