Live data from Hacker News

Amazon Glacier

aws.amazon.com

131–140 of 393 posts

Re: Amazon Glacier

#131
post #120

Amazon should complement this service with data contact centers which are connected to their data centers network. Then people could go to these centers in person and hand over their hard drives full of data for back up. It will be like bank lockers but only digital. At this low price people would want to upload terabytes of data which will be pain to upload/download.

You mean like http://aws.amazon.com/importexport/ ?

Re: Amazon Glacier

#132
post #121

Earlier quoted context omitted.

Yeah, I'm kind of wondering the same thing. It's certainly the kind of timeframe that changes your perspective. Maybe a tiny bit of Danny Hillis rubbed off on me from working at Applied Minds (man, I sure hope so!) Because as we answer issues of cost and availability, a logical thing to wonder is "how long can I really depend on it though?" As quickly as cloud services (where "lifetimes" are measured at six years) ha…

And, since you mentioned Danny Hills, it's also worth mentioning that (ironically?) Jeff Bezos is one of the principle supporters of http://longnow.org/clock/ .

And that makes me think of Anathem and the potential issues around long-term data storage that is capable of surviving through falls of civilisations and/or sacks of storage areas.

Re: Amazon Glacier

#133
post #90

I'm a long time user of backblaze, and I'm a big fan of the product - it does a great job of always making sure my working documents are backed up, particularly when I'm traveling overseas, and my laptop is more vulnerable to theft or damage. With that said - Backblaze is optimized for working documents - and the default "exclusion" list makes it clear they don't want to be backing up your "wab~,vmc,vhd,vo1,vo2,vsv,v…

Home-use is probably the only situation Glacier is good for backup though. A home user is fine with a 3.5-4 hour window before their backup becomes available for download (as it will probably take them days to download it anyway). In a corporate environment, I don't want to wait around for 3.5-4 hours before my data even becomes available for restore in a disaster recovery situation. Seems good for archive-only in a…

You can keep your recent backups in a fast access location for disaster recovery. It would be a good place to keep older backups though that you don't need to access in an emergency.

Re: Amazon Glacier

#134
post #69
post #41

This is a really good offering for media that you typically will keep locally for instant access, yet you want to have an off-site backup in a way that lives for a very long time. Dropbox should work here, but it's simply too expensive. My photo library is 175GB. That isn't excessive when considering I store the digital negatives and this represents over a decade. I don't mind not being able to access it for a few ho…

sigh Dropbox should not be used as a backup system. A system that synchronizes live should never get that role, except if they can guarantee that old data is never overwritten, new is always appended. This is not the case with dropbox - I've experienced multiple scary occurrences of old versions going to the nirvana with certain user actions. In certain cases the old data just appears to be gone, in other cases the w…

The ideal solution is the Quadfecta (is that a word?) - Dropbox for (in my experience) excellent versioning/synchronizing (Never failed me) + Backblaze (or its ilk) for continuous Offline backups + Super Duper (weekly/whenever) - for Image Backups - + Something (Arq?) on top of Glacier for long time off-site-archival.

For $50 in software (Arq+SuperDuper), $100 for an external HD, and less than $25/month ($4-backblaze, $10 Dropbox, $10 Glacier) you have a backup system that is next to air tight for a Terabyte of Data and a working set (on dropbox) of 100 Gigabytes.

Re: Amazon Glacier

#135
They should provide a "time capsule" option - pay X dollars, and after a set number of years, your data archive will be opened to the public for a given amount of time.

There'd be no better way to ensure that information would eventually be made public.

Re: Amazon Glacier

#136

Earlier quoted context omitted.

The dropbox client itself keeps the old versions of files around for a few days. It doesn't matter if the backend gets confused.

Yes, I'm aware of that. This helps, up to the point where you don't notice that something's gone for 'a few days'. Let's say you have your student project's folder shared with two other colleagues. You take a few days off, use your PC for casual browsing, meanwhile your colleagues are working on the project. At the end, one of them (who doesn't quite understand how dropbox works) deletes the files while the other is…

The colleague that was actively using the files doesn't notice them disappear?

Can you explain how dropbox has lost history for you in the past?

Re: Amazon Glacier

#137
post #27
post #21

Earlier quoted context omitted.

I think it means there is a small chance enough hard drives might fail at the same time that there happen to be no backups of those drives. They make so many backups so quickly that there is only a 0.00000000001% (I didn't count the zeros) chance of this occuring.

Which of course means that (if they're telling the truth) the probability of losing your data mostly comes from really big events: collapse of civilization, global thermonuclear war, Amazon being bought by some entity that just wants to melt its servers down for scrap, etc. (Whose probability is clearly a lot more than 10^-11 per year; the big bang was only on the order of 10^10 years ago.)

per object. So although the chance of losing any particular object is tiny, the chance of you losing something is proportional† to the number of objects. Still extremely small.

†roughly proportional if you have << 1e11 objects

Re: Amazon Glacier

#138
post #88

Earlier quoted context omitted.

No sigh was necessary, I understand. I even mentioned that I have a NAS and off-site backup... so either you didn't read it all or you stopped the moment you encountered the word "Dropbox" and started typing. That said, people DO use DropBox as backup. If you take a walk around the British Library and asked every PhD student working there how they "Backup" their research and work in progress, I bet every single perso…

"I even mentioned that I have a NAS and off-site backup... so either you didn't read it all or you stopped the moment you encountered the word "Dropbox" and started typing." No offense, but this post doesn't quite match what you wrote originally. My sigh was in response to the phrase "Dropbox should work here, ...". You didn't state security as your concern as to why not use Dropbox, rather it was cost. This might le…

"No offense, but this post doesn't quite match what you wrote originally. My sigh was in response to the phrase "Dropbox should work here, ...". You didn't state security as your concern as to why not use Dropbox, rather it was cost. This might lead someone who only needs A fair point.

In my case, Dropbox use is in addition to local RAID (scratch) NAS (network scratch, access of larger files) + off-site (backup).

I only use Dropbox for syncing and sharing.

The sync vs backup is an interesting one, simply because most consumers couldn't tell you the difference.

For example: Q: "Are your contacts backed up?". A: "Yes, they're sync'd to Google".

I did conflate my scenario with thinking about my girlfriend's peers in my post. And then reacted from my perspective again... my bad.

"Agreed. IMO no consumer backup system is quite there yet. TimeMachine is very close, if only it would do better logging and have some more intelligence about warning messages"

Vigorous agreement here too, except for the TimeMachine bit as that is Mac only and doesn't work for <insert any other system or device that isn't Apple Mac OSX).

Re: Amazon Glacier

#139
post #102
post #93

Earlier quoted context omitted.

1 x retrieval request per archive (it's really designed to store a small number of large files/tars) plus $0.12/GB... therefore, 4TB = $480.

No, that's not right. That's only the data transfer rate. If you check the FAQ, they bill you based on your peak hourly retrieval rate. If I download 4TB at say...10MB/s, not only do I need to pay $480, but I also have to pay ~$257 as a retrieval fee? Their wording is confusing, but ignoring the free retrieval amount (negligible difference on a 4TB transfer): Fee = Peak hourly retrieval * number of hours in month * $…

This is so confusing. So apparently if you spend the entire month retrieving the data at 1.6MB/s it only costs $40 plus transfer fees? And more importantly, how do you throttle your retrieval?

Edit: So I'm working through a scenario in my head and trying to figure out how charging based on the peak hour isn't completely ridiculous.

I have 8GB stored to try out the system. This costs a whopping dollar per year. One day I decide to test out the restore feature. So I go tell Amazon to get my files and wait a few hours. When Amazon is ready, I hit download. I'm on a relatively fast cable connection so the download finishes in an hour. I look at the data transfer prices and expect to be charged one dollar.

But I didn't take into account this 'peak hour' method. I just used roughly 8GB/hour over the minimal free retrieval. This gets multiplied out times 24 hours and 30 days to cost 8 * 720 * $0.01 = $57. Fifty-seven times my annual budget because I downloaded my data too quickly after waiting hours for Amazon to get ready.

Re: Amazon Glacier

#140
post #44

Earlier quoted context omitted.

Kinda: "In the coming months, Amazon S3 will introduce an option that will allow customers to seamlessly move data between Amazon S3 and Amazon Glacier based on data lifecycle policies."

URL? Can't find this on the website, blog or twitter account.

It's in the Glacier FAQ: http://aws.amazon.com/glacier/faqs/#How_should_I_choose_betw...
Post reply on HN