Live data from Hacker News

Internet Archive as a default host-of-record for startups

twitter.com

91–100 of 188 posts

Re: Internet Archive as a default host-of-record for startups

#91
post #87

The idea is more interesting when you think about scrapers. Take any ecommerce website, there are several scrapers that download all pages every hours, it would be more efficient if a provider had a live copy of the website and then serve the requests to the scrapers or could even send webhook. A website could handle tons of scrappers without having high bandwidth, only the provider will need high bandwidth. The issu…

(in the nicest way and since your post is recent, please edit s/scrappers/scrapers! It very much changed how I read it first time around -- I thought you were referring to a type of failed startup!)

Re: Internet Archive as a default host-of-record for startups

#92
post #27

Earlier quoted context omitted.

I agree, i'm a passionate photographer and i could pay good money to know that my pictures could be seen long time after my death. Maybe startups exists that do this, but they will die, i need something with enough critical mass that i can trust.

Perhaps have a look at Arweave ( https://www.arweave.org ).

That doesn't address the concern re: something needing critical mass to increase its chance of survival over a longer term.

Really the way I see it outside a few large banking firms, its kind of hard to be sure any provider of digital services would be around in the 50+ year term for this kind of public archive.

I hope the Internet Archive manages it.

EDIT: I do worry the IA has a bit of a lightning rod effect with skirting issues re: legality of archiving content. IMO its no guarantee it survives any significant time span either.

Re: Internet Archive as a default host-of-record for startups

#93
post #87

The idea is more interesting when you think about scrapers. Take any ecommerce website, there are several scrapers that download all pages every hours, it would be more efficient if a provider had a live copy of the website and then serve the requests to the scrapers or could even send webhook. A website could handle tons of scrappers without having high bandwidth, only the provider will need high bandwidth. The issu…

Is scraper one “p” or two? My inclination would be scrapper is “scrap”-er rather than “scrape”-er.

Re: Internet Archive as a default host-of-record for startups

#94
post #87

The idea is more interesting when you think about scrapers. Take any ecommerce website, there are several scrapers that download all pages every hours, it would be more efficient if a provider had a live copy of the website and then serve the requests to the scrapers or could even send webhook. A website could handle tons of scrappers without having high bandwidth, only the provider will need high bandwidth. The issu…

Scrapers. A scrAper is the thing you use to remove ice from windscreens or scrape data from websites. A scraPPer is someone who likes getting into fights or collecting scrap metal.

Re: Internet Archive as a default host-of-record for startups

#95
post #9

The feature I most want from the Internet Archive is the ability to donate them an old domain name and enough cash to renew it for the next hundred years such that they can keep an archived version of a site available (without breaking any incoming links) for a very long time. They would also need to be able to handle legal administration costs of things like DMCA take-down notices, but I assume they already have to…

> The feature I most want from the Internet Archive The feature I most want from IA is a streamlined system to delete content they have archived on domains that I own, including a proper privacy law compliance effort on their part. They have intentionally made it a difficult, manual process to get content removed. They operate as a de facto malicious crawler. They massively violate GDPR with how they operate and few…

> When IA has to comply with laws like GDPR, that's the end of IA.

Will you be happy when you've burned down that library?

Re: Internet Archive as a default host-of-record for startups

#97

I don't understand why Carmack thinks blockchain should be a component of this. Anyone care to elaborate on how that would make this easier/better?

.... He says it right after "to make internet applications that could outlive companies". If it's on the blockchain it doesn't matter if the company storing all of the archives shuts down, the content would still exist, forever, until their is a network running the chain. I suppose something like Torrent could be used?

Re: Internet Archive as a default host-of-record for startups

#99
post #59

Earlier quoted context omitted.

I think he's referring to something like IPFS. https://en.wikipedia.org/wiki/InterPlanetary_File_System http://ipfs.io You can put the storage costs on the nodes because storage at archive.org's scale adds up, especially when it's run by volunteers.

It looks like you could use IPFS to accomplish this without using a blockchain.

[deleted]

Re: Internet Archive as a default host-of-record for startups

#100

I think it's interesting to think about what we have lost because we couldn't keep everything from a 100 years ago and what society 100 years from now will be grateful we preserved. Off the top of my head, we lost a lot of common wisdom in dealing with the flu pandemic of 1918 because personal letters and most newspapers were not preserved. I think 100 years from now they might wish we had preserved more from margina…

Hard to know what will be of interest for future historians. Some things in which we place great value can be considered irrelevant, while some of our junk can become historical gold.

The mundane of today is very insightful for tomorrow's historians.

Its fascinating when you start looking into any historical time period (you wouldn't even need to go far back), before a lot of details are educated guesses. Since no one chose to record the mundane in detail or it failed to preserve over time.

Post reply on HN