Live data from Hacker News

Thank you for helping us increase our bandwidth

blog.archive.org

171–180 of 207 posts

Re: Thank you for helping us increase our bandwidth

#171

Earlier quoted context omitted.

Aren't the costs of getting that data out of the backup much larger than the cost of keeping it in the first place, to the point that when you actually need to restore a large backup, it turns out it was better to have been managing it yourself? That's the impression I got from various HN comments on the topic over the years.

The cost is only there if you transfer out of aws. Something like glacier will have a retrieval time on the order of hours or days.

Glacier Deep Archive does charge for retrievals at $0.02 per GB and additional $0.01 per 1000 such requests (both of which are $0.00 for Standard S3). PUT, LIST, DELETE are at $0.05 per 1000 requests, 10x the Standard S3 rates.

https://aws.amazon.com/s3/pricing/

Re: Thank you for helping us increase our bandwidth

#172
post #100

Earlier quoted context omitted.

BackBlaze B2 is $5/TB/month Azure Archive is $2/TB/month ($1.68 if reserved) AWS Glacier Deep Archive is $1/TB/month GCP Cloud Storage Archive is $1.20/TB/month Of course, there can be i/o and network charges, and different levels of redundancy (but possibly bulk discounts)...but the bare storage costs for for 50 PB per year would be roughly $600k - $3 MM/y.

Aren't the costs of getting that data out of the backup much larger than the cost of keeping it in the first place, to the point that when you actually need to restore a large backup, it turns out it was better to have been managing it yourself? That's the impression I got from various HN comments on the topic over the years.

AWS has Snowball and Snowmobile though only used former to reduce data transfer costs. Dont remember what other savings are in there. Like is there price reduction if use with Glacier or not.

Re: Thank you for helping us increase our bandwidth

#173
post #81

Archive.org works surprisingly well as a general purpose web proxy. Just prefix the URL, e.g., http://example.com , with https://web.archive.org/save/ , e.g., https://web.archive.org/save/http://example.com The aesthetic intrusiveness of the archive.org header and footer are minimal since I use a text-only browser that has no Javascript engine. Sometimes I get "This url is not available on the live web or can not be…

I use a custom browser keyword search to find existing archived pages before saving one, personally. Eg: ar for: https://wayback.archive.org/web/*/%S I'd imagine it would be useful for IA to implement some message for scenarios where a page has already been saved within a certain timespan and provide both a link to the already saved version and offer to save again. As this would mitigate mass savings of an identical…

That's a good point. I mainly use it for browsing websites that change daily as well as ones with many successive pages, e.g., 1, 2, 3, etc. that IA does automatically not crawl. Thus, not many existing copies if any.

HAproxy changes the Host header and modifies the URL. I can either use the text-only browser's http-proxy option or I can direct the request to the web.archive.org backend by adding a custom HTTP header to the request. If I am not mistaken, the so-called "modern" browsers do not have built-in capability to add headers.

Re: Thank you for helping us increase our bandwidth

#174

archive.org feels like an irreplaceable treasure, the Wayback Machine alone is a time capsule of our digital history. I donate to them monthly and know a lot of other people do as well, so I don't worry much about their financial stability. I'm more worried about external pressures taking content down. I hope the data is backed up six ways to sunday, and that somewhere there's a plan to make it all accessible if Inte…

They could set up a torrent that people could download parts of, I think it's an ideal system for a distributed backup.

But it probably needs tweaking for that purpose; for one you need to ensure that the data is evenly distributed, and second you're dealing with data that is appended to regularly.

But I think it can be done.

Re: Thank you for helping us increase our bandwidth

#175

Earlier quoted context omitted.

Aren't the costs of getting that data out of the backup much larger than the cost of keeping it in the first place, to the point that when you actually need to restore a large backup, it turns out it was better to have been managing it yourself? That's the impression I got from various HN comments on the topic over the years.

AWS has Snowball and Snowmobile though only used former to reduce data transfer costs. Dont remember what other savings are in there. Like is there price reduction if use with Glacier or not.

Isn't that inbound only? Getting the data out again is also required.

Re: Thank you for helping us increase our bandwidth

#177

Earlier quoted context omitted.

AWS has Snowball and Snowmobile though only used former to reduce data transfer costs. Dont remember what other savings are in there. Like is there price reduction if use with Glacier or not.

Isn't that inbound only? Getting the data out again is also required.

You can export data via Snowball as well.

https://docs.aws.amazon.com/snowball/latest/ug/create-export...

Re: Thank you for helping us increase our bandwidth

#178

archive.org feels like an irreplaceable treasure, the Wayback Machine alone is a time capsule of our digital history. I donate to them monthly and know a lot of other people do as well, so I don't worry much about their financial stability. I'm more worried about external pressures taking content down. I hope the data is backed up six ways to sunday, and that somewhere there's a plan to make it all accessible if Inte…

I've thought for a long time why I absolutely agree with this sentiment and the best I can come up with is that the Internet Archive feels to me like the embodiment of the old internet, not this marketing driven, data steal-and-sell, VC backed cyberpunk dystopia Internet that we have everywhere else.

It's the same sort of feeling as I why I enjoy this site, wikipedia, and so on. Bonus, IA is also a bit of an organizational mess, but rewards the adventurer with rich treasure.

(I was recently on a Korean history kick and came across not only 1, but several entirely different first hand books written by visitors in the late 19th and early 20th century, scanned in, freely accessible/downloadable, in a variety of formats, and with an excellent on-web reader. These books are so out of print I checked with three local counties for copies and none of them even have references to any of them in their catalog -- treasure!)

Re: Thank you for helping us increase our bandwidth

#179
post #100

Earlier quoted context omitted.

BackBlaze B2 is $5/TB/month Azure Archive is $2/TB/month ($1.68 if reserved) AWS Glacier Deep Archive is $1/TB/month GCP Cloud Storage Archive is $1.20/TB/month Of course, there can be i/o and network charges, and different levels of redundancy (but possibly bulk discounts)...but the bare storage costs for for 50 PB per year would be roughly $600k - $3 MM/y.

Aren't the costs of getting that data out of the backup much larger than the cost of keeping it in the first place, to the point that when you actually need to restore a large backup, it turns out it was better to have been managing it yourself? That's the impression I got from various HN comments on the topic over the years.

Low cost to insert, low cost to keep it there, high cost to retrieve is exactly the combination you want when looking at disaster backup solutions, since you don't intend to retrieve the data frequently. Buy some earthquake insurance (I know, easier said than done) and only pay for 1/20 of the retrieval cost.

Re: Thank you for helping us increase our bandwidth

#180
post #61

archive.org feels like an irreplaceable treasure, the Wayback Machine alone is a time capsule of our digital history. I donate to them monthly and know a lot of other people do as well, so I don't worry much about their financial stability. I'm more worried about external pressures taking content down. I hope the data is backed up six ways to sunday, and that somewhere there's a plan to make it all accessible if Inte…

I'm very worried about its backups and what happens when the next Big One hits SF. As far as I've ever been able to determine from talking to anyone at IA (e.g., Kahle, Scott), they don't really have any sort of backups that could actually be restored from in a disaster situation.

Whatever happened to their copy in Canada?

https://blog.archive.org/2016/12/03/faqs-about-the-internet-...

Post reply on HN