Live data from Hacker News

Backblaze Durability Is Eleven 9s – And Why It Doesn’t Matter

backblaze.com

61–70 of 190 posts

Re: Backblaze Durability Is Eleven 9s – And Why It Doesn’t Matter

#61
The best durability is probably achieved by Amplidata, but it does not matter.

You need to do the same calculation for your meta data, which is probably not erasure coded. If you lose this, you don't lose your data, but you no longer know where you put it.

So you probably add your meta data to your data as well in some kind of recoverable format. That's fine, it means that you can harvest the meta data again.

But how long does this take ?

Re: Backblaze Durability Is Eleven 9s – And Why It Doesn’t Matter

#62
post #7
post #2

This was an interesting read, both the points made about durability, as well as the in-depth math. However, what stood out to me most was the line: Because at these probability levels, it’s far more likely that: - An armed conflict takes out data center(s). - Earthquakes / floods / pests / or other events known as “Acts of God” destroy multiple data centers. - There’s a prolonged billing problem and your account data…

I've posted the math here before but if we assume that an asteroid hits the earth every 65 million years and wipes out the dominant life forms, then this fact alone puts your yearly durability at a maximum of ~8 nines. The point about billing is better, though. My other concern is that a software bug, operator error, or malicious operator deletes your data.

That's why no sane entity would use one earth. Use two and you can quickly recover.

Re: Backblaze Durability Is Eleven 9s – And Why It Doesn’t Matter

#63

This analysis is simplistic. Correlated failures are common in drives. That could be a power surge taking out a whole rack, a firmware bug in the drives making them stop working in the year 2038, an errant software engineer reformatting the wrong thing, etc. When calculating your chance of failure, you have to include that, or your result is bogus. Eg. Model A of drive has a failure rate of 1% per year, but when fail…

Nice analysis.

Re: Backblaze Durability Is Eleven 9s – And Why It Doesn’t Matter

#64
post #5

This was an interesting read from a technical point of view, but also well written and refreshingly transparent. I found the discussion about why it doesn't matter when you start talking about 11 nines of reliability to be hilariously true. At the end of the day we're still flawed humans living in a hostile universe, and no matter how foolproof we make the technology, there are some weaknesses that just can't be elim…

It's 'foolproof', like dog proof but for fools.

Hah, good catch, I know that but somehow made the error anyway. I corrected it.

Re: Backblaze Durability Is Eleven 9s – And Why It Doesn’t Matter

#65
post #2

This was an interesting read, both the points made about durability, as well as the in-depth math. However, what stood out to me most was the line: Because at these probability levels, it’s far more likely that: - An armed conflict takes out data center(s). - Earthquakes / floods / pests / or other events known as “Acts of God” destroy multiple data centers. - There’s a prolonged billing problem and your account data…

For a consumer the cheapest and easiest way to backup important documents or files is to encrypt it and store it across multiple storage providers, e.g. Dropbox and Google Drive.

They usually give you a reasonable amount of free storage, and it's unlikely all accounts would be terminated or locked at the same time.

And of course, you should always have your local backups as well.

Re: Backblaze Durability Is Eleven 9s – And Why It Doesn’t Matter

#66

Earlier quoted context omitted.

That could be a power surge taking out a whole rack This failure mode, at least, is already accounted for by sharding data across cabinets: Each file is stored as 20 shards: 17 data shards and 3 parity shards. Because those shards are distributed across 20 storage pods in 20 cabinets, the Vault is resilient to the failure of a storage pod, or even a power loss to an entire cabinet. https://www.backblaze.com/blog/vaul…

I thought AZs were in the same physical location, just separate networks, no?

Others have answered, but I think the general principle for a Region is that between the AZs you have less than 1ms latency, physical separation between the data centers, but they may still be in the same floodplain, be able to be hit by the same hurricane, etc.

More info here if you want to see how regions are structured at a high level of detail: https://youtu.be/AyOAjFNPAbA?t=15m21s

Re: Backblaze Durability Is Eleven 9s – And Why It Doesn’t Matter

#67
post #7
post #2

This was an interesting read, both the points made about durability, as well as the in-depth math. However, what stood out to me most was the line: Because at these probability levels, it’s far more likely that: - An armed conflict takes out data center(s). - Earthquakes / floods / pests / or other events known as “Acts of God” destroy multiple data centers. - There’s a prolonged billing problem and your account data…

I've posted the math here before but if we assume that an asteroid hits the earth every 65 million years and wipes out the dominant life forms, then this fact alone puts your yearly durability at a maximum of ~8 nines. The point about billing is better, though. My other concern is that a software bug, operator error, or malicious operator deletes your data.

Isn't the expected life span of the company more limiting? Plenty of cloud storage companies go out of business (typically, they run out of money). You can apply Gott's law to this. It's pretty grim

Re: Backblaze Durability Is Eleven 9s – And Why It Doesn’t Matter

#68
post #18

Financial failure or service shutdown by the provider is the highest risk for long term storage. The backup services CrashPlan, Dell DataSafe, Symantec, Ubuntu One, and Nirvanix all shut down. Nirvanix only gave two weeks notice for users to save their data.[1] [1] https://www.computerweekly.com/opinion/Nirvanix-failure-a-bl...

add Bitcasa to the list: https://en.wikipedia.org/wiki/Bitcasa

Re: Backblaze Durability Is Eleven 9s – And Why It Doesn’t Matter

#69

Earlier quoted context omitted.

>The chance they can recover all their data is only (1-0.003)^100000... Which means every customer suffers data loss :-( You just made the same mistake you're criticizing. You assumed the 100K files were uniformly and independently spread. They're also likely clustered, and perhaps not even at the same data center. Given the variety of drives Backblaze uses, the drives are also not likely to all be the same model, so…

>You just made the same mistake you're criticizing Yes. Even if you take it into account, the vast majority of customers will see data loss, assuming random (but not even) shard distribution. I rounded and approximated to avoid explaining too many probability rules... As you can see, it gives the same result at the end.

>it gives the same result at the end

No, it doesn't. You picked 3% of all drives failing out of the blue. Your next estimate was an order of magnitude too high. Your last assumption of p^# drives is not reasonable.

The proof is in reality. Backblaze has run over a decade, with all sorts of hardware failures, server configs, running many drive models through their lifetime, across manufacturers, across technologies, across multiple datacenters, and had not seen the level of failures you claim they will.

So I suspect their method of estimating is more accurate than yours. So far it matches reality much better.

Re: Backblaze Durability Is Eleven 9s – And Why It Doesn’t Matter

#70
post #8

It still wouldn't upload my 1TB in back-ups in an entire month. Amazon Drive back-up completed in 3 days. Their pricing is amazing, but saving money on a back-up solution that doesn't seem as good as the other cloud storage providers is a dangerous game.

I use B2 and I have no problems uploading large files quickly. Is this for the consumer option?
Post reply on HN