Live data from Hacker News

Backblaze Hard Drive Stats for 2018

backblaze.com

121–130 of 165 posts

Re: Backblaze Hard Drive Stats for 2018

#121
post #77

Earlier quoted context omitted.

Disclaimer: I work at Backblaze so I'm biased. :-) > If you have terabytes to back up, are there still any backup services left that'll let you ship them a drive for a faster initial backup? Backblaze offers a "Backblaze Rapid Ingest Fireball" to allow you to ship us 60 TBytes of data on an appliance. https://www.backblaze.com/blog/introducing-backblazes-rapid-... If you only have 2 - 10 TBytes, I suggest you get a f…

Hi Brian, Thank you for reaching out. I love it when company folks will get into the conversation on HN. As far as connection, I've got 300/300 from Frontier. It's good, and I can help out my friends. But more questions below... I had been using CrashPlan for years. Converted to their business plan when they decided to ditch the Consumer stuff. My confidence in their viability/user experience has eroded. I have perso…

Disclaimer: I work at Backblaze so you should always view my answers skeptically. :-)

> Color me skeptical that folks outside of the public cloud providers are going to be around in 10 years.

Backblaze is now 12 years old, and we're actually kind of unique in that we have never raised any significant VC funding and we're (slightly) profitable. We have run the business entirely as a business (not on unsustainable VC dollars), and we aren't planning on going anywhere. Backblaze is employee owned and run, the only voting board members are the original five founders. We have a couple of "board observers" from outside the company for "adult supervision and experienced advice", but they cannot control us, and they cannot even vote.

Side Note: The Backblaze founders and a good portion of the staff all came from the same previous startup/company, and in that case the VCs forced us to sell it, which murdered it. The whole reason we self funded Backblaze and ran a sustainable (profitable) business was a reaction to how horrible that situation was.

Backblaze currently has 803 PBytes of storage in our three datacenters, and business is really going well and we are growing quite healthy. We also understand (and talk about among ourselves) the large responsibility here, which is realistically that amount of data cannot ever be moved. If Backblaze decided on a whim to shut down, we would seriously, SERIOUSLY hurt or even manage to destroy thousands of businesses which depend on us. So we are not going to do that.

Re: Backblaze Hard Drive Stats for 2018

#122

Curious if anyone knows how they calculate drive days. For example the first drive on their report (hgst 4tb) has a count of 50 but total days of 23069. If I take the 50 by the days it should be 18250, so not sure where the extra 4k in days is coming from. Retired drives or something?

I think it is over the time they have had the drive, not over the reporting interval. It's a measure of the age of the drives. Assuming their drives are operating 24/7, that means that those particular 50 drives have been in service an average of 461 days.

I'd expect on next year's report, those particular drives will show up as 49 drives with around 42000 drive days, assuming they aren't replaced by then.

Re: Backblaze Hard Drive Stats for 2018

#123

Thank you Backblaze! I love your reports. What is your procedure/policy on which disks to use in the pods? Do you try and maybe control the risk by using different harddisk brands in a single storage pod? Or do you just not care, because there have never been 3 pods dead at the same time? :) Do you still use 17 data plus 3 parity shards?

Their Q3 2018 stats had a bit of info on the lifecycle of introducing new disks:

https://www.backblaze.com/blog/2018-hard-drive-failure-rates...

> In Q3 we added 79 HGST 12TB drives (model: HUH721212ALN604) to the farm. While 79 may seem like an unusual number of drives to add, it represents “stage 2” of our drive testing process. Stage 1 uses 20 drives, the number of hard drives in one Backblaze Vault tome. That is, there are are 20 Storage Pods in a Backblaze Vault, and there is one “test” drive in each Storage Pod. This allows us to compare the performance, etc., of the test tome to the remaining 59 production tomes (which are running already-qualified drives). There are 60 tomes in each Backblaze Vault. In stage 2, we fill an entire Storage Pod with the test drives, adding 59 test drives to the one currently being tested in one of the 20 Storage Pods in a Backblaze Vault.

Re: Backblaze Hard Drive Stats for 2018

#124
post #53

Earlier quoted context omitted.

Disclaimer: I work at Backblaze so I'm biased. :-) > If you have terabytes to back up, are there still any backup services left that'll let you ship them a drive for a faster initial backup? Backblaze offers a "Backblaze Rapid Ingest Fireball" to allow you to ship us 60 TBytes of data on an appliance. https://www.backblaze.com/blog/introducing-backblazes-rapid-... If you only have 2 - 10 TBytes, I suggest you get a f…

but don't most ISPs have datacaps nowadays, Comcast has 1TB

> Comcast has 1TB/month data cap

I have three suggestions, but I have only tried the second suggestion, so please do your own research before my bad advice costs you a lot of money. :-)

1) Comcast (at least in many places) allows you to exceed your bandwidth cap for two months before clamping down on you. I think they are trying to prevent serious long term abuse, not a one time overage. So if you have 3 TBytes and can get it uploaded in one or two months, just do it, apologize, and it won't cost you anything. Backblaze only does "incrementals" after the initial upload.

2) Personally I have Comcast and I pay them an extra $30 or so per month for "unlimited" (remove the cap). Now when I look at my usage, my family stays just under the 1 TByte limit ANYWAY, so this is wasted money, but I don't want to stress about it, and I run like 5 Nestcams CONSTANTLY streaming video, plus my family loves Netflix, so I just drop the $30 and relax. So you could call up Comcast and change over to "unlimited" if you can afford it.

3) A modification of #2 that I have NOT TRIED is to raise it to "unlimited" for the duration of the initial backup. I don't know if you have to commit to a year of unlimited bandwidth, or six months, or if you can change at any time?

Re: Backblaze Hard Drive Stats for 2018

#125

Thank you Backblaze! I love your reports. What is your procedure/policy on which disks to use in the pods? Do you try and maybe control the risk by using different harddisk brands in a single storage pod? Or do you just not care, because there have never been 3 pods dead at the same time? :) Do you still use 17 data plus 3 parity shards?

[deleted]

Re: Backblaze Hard Drive Stats for 2018

#126

Earlier quoted context omitted.

Disclaimer: I work at Backblaze. > When a drive fails, it's effectively a brick with no terabytes. Interesting factoid: that isn't always true. What you describe is actually the CLEANEST type of failure, the drive suddenly becomes a brick. We replace the drive and rebuild it from parity. A way more interesting failure is when disk blocks start going bad at an unacceptable rate. Backblaze splits your data across 20 di…

(thanks a lot - all extremely interesting) >> We can't have more than 3 drives fail in any one group of 20 drives... Wow, for me, subjectively, an low threshold - and I underderstand that each drive being hosted on a different machine protects you as well from a machine/controller failure (happened to me twice with the controller - both times it was very hard to diagnose and the experience in general has been terribl…

> Do you have as well "backups"? Or is that in the hands of the customers/users?

If you store data in Backblaze, there is no "backup" of that data. If Backblaze ever lost 4 drives simultaneously and could not recover the data, the customer would lose data. This is much like Amazon S3.

In general, we recommend a 3-2-1 backup strategy where there are 3 copies of the data, at least 2 copies on your site, and 1 copy in the cloud. You can read about that philosophy in our blog post here: https://www.backblaze.com/blog/the-3-2-1-backup-strategy/

Re: Backblaze Hard Drive Stats for 2018

#127

Earlier quoted context omitted.

(thanks a lot - all extremely interesting) >> We can't have more than 3 drives fail in any one group of 20 drives... Wow, for me, subjectively, an low threshold - and I underderstand that each drive being hosted on a different machine protects you as well from a machine/controller failure (happened to me twice with the controller - both times it was very hard to diagnose and the experience in general has been terribl…

> Do you have as well "backups"? Or is that in the hands of the customers/users? If you store data in Backblaze, there is no "backup" of that data. If Backblaze ever lost 4 drives simultaneously and could not recover the data, the customer would lose data. This is much like Amazon S3. In general, we recommend a 3-2-1 backup strategy where there are 3 copies of the data, at least 2 copies on your site, and 1 copy in t…

Thank you!

To summarize I understand: A) the local working copy (locally replicated in your case), B) the local backup and C) the cloud/very remote backup. B & C cover each other if any datacenter is completely wiped out.

Re: Backblaze Hard Drive Stats for 2018

#128

Earlier quoted context omitted.

Disclaimer: I work at Backblaze. > When a drive fails, it's effectively a brick with no terabytes. Interesting factoid: that isn't always true. What you describe is actually the CLEANEST type of failure, the drive suddenly becomes a brick. We replace the drive and rebuild it from parity. A way more interesting failure is when disk blocks start going bad at an unacceptable rate. Backblaze splits your data across 20 di…

>>So when a drive is HALF-FAILED, we even have a procedure to pull the drive out, and then opportunistically copy whatever files we can recover onto a new drive, then put the new drive back into production. Do I understand correctly, that when the drive is half-failed, you don't just say "it will probably completely stop working in the near future" and discard/replace it but keep using it?

If it was half failed we would DEFINITELY pull the drive out because it is already 1 drive down out of 20 for half the files. A lot of times the IT guys will make a judgement call that a drive is acting funny or slightly off so they just "fail it on purpose" which means yank it and replace with a new drive. We have done this just because a drive is "slow" (slow can mean the drive is having trouble writing data reliably on one attempt), or because some SMART stat looks wonky.

To provide more color, if a 20 drive "tome" (as we call it) is 1 drive down, we don't even wake people up in the middle of the night, but Backblaze datacenter employees replace it when they arrive at the datacenter the next day at 8am. All drives having problems are replaced by 5pm when the employees go home. This is completely business as usual, about 5 - 10 drives fail every day.

However, if 2 drives fail out of 20 (or 1.5 in our example above), pagers go off, people wake up and get out of bed at 3am and start driving towards the datacenter. Or we employ "remote hands" to swap the drives immediately, it depends on the capabilities of the night crew in the datacenter which varies by datacenter. "remote hands" is a contract service where semi-skilled technicians work for the datacenter and we can pay them $80/incident or there abouts to do things you can only do "in person" like replace drives. All the pods (where data is stored) have "base board management" which means as long as they are powered up and online we can log in remotely from home or office to figure out what is going on and fix a variety of problems. AUTOMATICALLY if 2 drive fail we stop sending any data into that "tome" of 20 drives. We have found that writing to drives causes more failures, so not writing to them is safer.

If 3 drives fail, it is instantly a "Red Alert" at Backblaze and a whole lot of official procedures kick in. An "incident manager" is assigned and the whole company's number one concern is to drop EVERYTHING and never sleep again until the Red Alert is lowered to Yellow. We light up a "situation room" (in Slack - our internal chat tool) and information and status is relayed through that.

SIDE NOTE: Backblaze has a relationship with an excellent company named "DriveSavers" who can recover SOME data off of failed drives. This is very expensive (thousands of dollars per drive) so we only do it to test the procedure and then in extreme situations. Three drives down is an extreme situation and extremely rare, so ALL OF THE THREE FAILED DRIVES would be immediately hand carried to DriveSavers even while we rebuild the customer data from parity. Notice Backblaze STILL has a complete copy of the customer data on 17 drives -> But if a 4th drive dies, the hope is we can recover at least one of the drives via DriveSavers thus saving the customer data. (We need at least 17 out of 20 drives in a "tome" to reconstruct the data.) In our experiments, DriveSavers seems to recover about half the drives, or in some situations half the data from a drive (imagine if 1 platter on a drive has a head crash and is destroyed, but the other platters are fine). We have made the decision that it is less expensive (for the same durability) to pay DriveSavers the thousands of dollars rarely instead of increasing parity to allow reconstructing data from 16 out of 20 drives instead of the current 17 out of 20 drives.

Re: Backblaze Hard Drive Stats for 2018

#129

Earlier quoted context omitted.

> Do you have as well "backups"? Or is that in the hands of the customers/users? If you store data in Backblaze, there is no "backup" of that data. If Backblaze ever lost 4 drives simultaneously and could not recover the data, the customer would lose data. This is much like Amazon S3. In general, we recommend a 3-2-1 backup strategy where there are 3 copies of the data, at least 2 copies on your site, and 1 copy in t…

Thank you! To summarize I understand: A) the local working copy (locally replicated in your case), B) the local backup and C) the cloud/very remote backup. B & C cover each other if any datacenter is completely wiped out.

Correct.

> if any datacenter is completely wiped out

Correct. When all of our datacenters were in Sacramento, California, some customers told us they were concerned because they were ALSO in Sacramento and a meteor could wipe out both their computer, the local backup, and Backblaze's cloud backup, all in one meteor strike.

While by default we put your data where it is convenient for Backblaze, we CAN work with customers (and have done so) to place their data in our Phoenix Arizona datacenter or one of our Sacramento datacenters if it is important. As we add our European region (coming soon) this will become a pull down menu for all customers. For now, we only work with larger customers to make sure the customer data lands in the correct location for them.

Re: Backblaze Hard Drive Stats for 2018

#130

Earlier quoted context omitted.

If you've never found anything with it, why do you keep doing it (see: the definition of insanity)?

Not the parent, but probably because 1. It's notorious that hard drives have a higher failure rate at the beginning of their lives than in the middle (see bathtub curve [0]). So it's not absurd to test them hard early on before writing any useful data and to do an early RMA. 2. The failure rate on drives is low enough that his methodology may be right but he still never has any failure in his life. Doesn't it make in…

It would depend on the effort to do the methodology vs. the expected return (savings of finding a failed drive times the probability).
Post reply on HN