Live data from Hacker News

Backblaze Durability Is Eleven 9s – And Why It Doesn’t Matter

backblaze.com

121–130 of 190 posts

Re: Backblaze Durability Is Eleven 9s – And Why It Doesn’t Matter

#121
post #40
post #31

Earlier quoted context omitted.

Also, there are only 3 hard drive manufacturers left. If one of them have a bug that affects across their product line, that can take out 1/3 of all hard drives.

It's vanishingly unlikely that a bug would affect all their drives (across all recent models, and only after burn-in) simultaneously, unless the drives are managed by a remote server with a SPOF.

You never experienced Death Star.

Re: Backblaze Durability Is Eleven 9s – And Why It Doesn’t Matter

#122
post #117
post #56

Earlier quoted context omitted.

That's one of the reasons to encrypt locally and only store encrypted data on backup services.

If Google suspends your account for background music playing in a YouTube video, you might still lose access to your files in Google drive / cloud - even if the files are encrypted.

You should compartmentalize those.

Re: Backblaze Durability Is Eleven 9s – And Why It Doesn’t Matter

#123
post #73

Earlier quoted context omitted.

A better example would be the ceramic bearing fiasco that NetApp experienced with Seagate. Seagate had switched to a floating ceramic bearing on one family of their fiber channel drives. In those drives one or more of the bearings would shatter and start spreading ceramic dust across the disk surface. This happened between 3 and 4 years of run time and the disk would rapidly fail after that happened. People that boug…

There was also a failure mode in Seagate drives where the bearing increased in stiction. As long as it was spinning, there was no problem. But if you spun it down, it might not spin up again. If you had a group of disks powered up for a long time, many could fail together at the next power cycle. A chaos monkey that randomly powers down disks one at a time can prevent this.

Though that sounds like a fixable failure mode, so it's more annoyance than tragedy.

Re: Backblaze Durability Is Eleven 9s – And Why It Doesn’t Matter

#124

Earlier quoted context omitted.

I was going to post essentially the same thing, so here is an upvote :-) While I always find storage analysis interesting (I spent 5 years at NetApp where it was sort of a religion :-)) some of the assumptions that Brian was tossing out are not good ones to make. (like the lack of correlation, or that Drive Savers will exist as a company 10 years from now). Still it does help you to understand the they take data avai…

Disclaimer: I'm the author of the blog post. :-) > some of the assumptions that Brian was tossing out are not good ones to make. We COMPLETELY welcome other analysis and listing other assumptions. Internally, we argued endlessly about why this or that wasn't totally accurate, and finally decided to publish the math WITH all of our assumptions exposed so you could be the judge. If Amazon wants to publish their assumpt…

> If Amazon wants to publish their assumptions for S3 for comparison, we're all ears.

The 2010 S3 calculation is obviously also wrong! I totally feel your pain in terms of wanting to have a directly comparable answer, but the reasons you give that "it doesn't matter" (and others, including correlated faults, software bugs, model risk, and security breaches) are actually reasons why the stated durability number is wrong. Honestly IMO a probability of 1-10^-11 is the wrong answer to pretty much any question; model risk is going to dominate that for any problem more complex than 1+1=2.

That said, although neither your system nor Amazon's should be expected to have anywhere near "eleven nines" of durability in reality, if as I understand it S3 is split across availability zones in a region and your product has all splits in a monolithic DC, I would expect S3 to come out ahead in a more careful analysis. (But note that S3 is not really a seamless multi region product, though there is an option to set up cross region replication.)

Re: Backblaze Durability Is Eleven 9s – And Why It Doesn’t Matter

#125
post #7

Earlier quoted context omitted.

I've posted the math here before but if we assume that an asteroid hits the earth every 65 million years and wipes out the dominant life forms, then this fact alone puts your yearly durability at a maximum of ~8 nines. The point about billing is better, though. My other concern is that a software bug, operator error, or malicious operator deletes your data.

Isn't the expected life span of the company more limiting? Plenty of cloud storage companies go out of business (typically, they run out of money). You can apply Gott's law to this. It's pretty grim

I guess this is what you're referring to: https://en.wikipedia.org/wiki/J._Richard_Gott#Copernicus_met...

Re: Backblaze Durability Is Eleven 9s – And Why It Doesn’t Matter

#126

I really want to like Backblaze and they seem to do a lot of good work, but whenever this comes up, I also feel responsible to let people know the dark side so they're informed at least. I've written in more detail before[0], but just to share the gotchas in case anyone here is thinking of switching to Backblaze: 1. They backup almost no file metadata. 2. The client is very slow (days or more) to add new files and th…

I didn't like their main client very much (though it was a while ago) but I'm still planning to use B2 for some things. So it depends on what you're buying from them.

Re: Backblaze Durability Is Eleven 9s – And Why It Doesn’t Matter

#127

Earlier quoted context omitted.

If it's your data center, you can plan the location to prevent most of this. There are locations where the natural hazards can be completely managed. (No tornados, fires, tsunamis, earthquakes, ...) So the power outage is the most likely thing to happen.

I guess. I feel like "a tornado can never happen" is a lot like those lines in the logs like "error: can't happen". It can't happen, but it does.

Theoretically, yes, it can happen.

Realistically, the chance of a tornado taking out the Swedish datacenter built inside a former nuclear bunker under 100ft of granite bedrock is so small that it probably doesn't affect the number of 9's that you can claim.

Re: Backblaze Durability Is Eleven 9s – And Why It Doesn’t Matter

#128
post #25

Earlier quoted context omitted.

> In that case, assuming 10,000 drives, to recover one dead 12TB drive and a recovery rate of even 10 MB/secs per machine, recovery of one drive should be done in under a second. I want to know where you can find a drive that can write 12TB/sec of data! (In other words, you clearly missed half the problem. To add a new replacement drive, you have to be able to write to it the data from an original drive. Also RS code…

Additionally, it implies that data needs to be spread across 10,000 drives, which is unrealistic anyway.

As to the one second claim, I think that's a math error, because even 10k * 10MB is only 100GB.

But spreading the data over 10k drives isn't unrealistic, it's a different architecture. Pick a different 20 drives for each file.

Working it through: Assume 200 machines with 50 drives each. Each machine has to read 1TB, transmit it over the network, do a parity calculation, and write out ~50GB. With dual 10gbps ports the bottleneck is the network, and if we dedicate one on each machine to the rebuild we get a 15 minute clock.

Not that having such a monolithic architecture is worth the complication and extra bugs.

Re: Backblaze Durability Is Eleven 9s – And Why It Doesn’t Matter

#129
post #127

Earlier quoted context omitted.

I guess. I feel like "a tornado can never happen" is a lot like those lines in the logs like "error: can't happen". It can't happen, but it does.

Theoretically, yes, it can happen. Realistically, the chance of a tornado taking out the Swedish datacenter built inside a former nuclear bunker under 100ft of granite bedrock is so small that it probably doesn't affect the number of 9's that you can claim.

Sadly not even Swedish nuclear bunkers are safe from disaster: https://www.theguardian.com/environment/2017/may/19/arctic-s...

> Arctic stronghold of world’s seeds flooded after permafrost melts > It was designed as an impregnable deep-freeze to protect the world’s most precious seeds from any global disaster and ensure humanity’s food supply forever. But the Global Seed Vault, buried in a mountain deep inside the Arctic circle, has been breached after global warming produced extraordinary temperatures over the winter, sending meltwater gushing into the entrance tunnel.

There are always unforeseen and unforeseeable risks associated with any location. You can mitigate them but you can't claim X number of 9s for a single physical datacenter.

Re: Backblaze Durability Is Eleven 9s – And Why It Doesn’t Matter

#130

> if you store 1 million objects in B2 for 10 million years, you would expect to lose 1 file. Can this be reformulated: you store 10 trln objects (e.g. 100TB of 10 byte records), you lose 1 record each year. Also curious what are the stats from other providers.

The raw number would imply that, but I'm pretty sure the math breaks down when you're storing 10 byte records.

Chunks of your data are going to be stored together, so it's a very small chance of losing a big block of 10 byte files. There's no failure mode that loses just one, and does so often.

Post reply on HN