Live data from Hacker News

Hard Drive Reliability Update – Sep 2014

backblaze.com

71–80 of 168 posts

Re: Hard Drive Reliability Update – Sep 2014

#71
post #44
post #29

Earlier quoted context omitted.

Do you have any data to back up any of these claims? Because now I have a problem: should I believe guys who have been running 38 petabytes of storage for several years now and regularly present the data they gathered, or should I believe you?

I've read his blog posts before. He's a good author. I don't believe anything factual/technical in my post disagrees with reality or any of his historical posts. I am open to the idea I'm misinterpreting how he presents enterprise vs consumer firmware loads. I assume its open knowledge that their secret sauce is using consumer hardware so assumptions about 15K fibre channel isn't relevant.

> Do you have any data to back up any of these claims?

Can you respond to this directly, please?

Re: Hard Drive Reliability Update – Sep 2014

#72

I love reading these posts from Backblaze, but what I never understand is that they are getting a cost of U$ ~0.05/GB with their storage pods: https://www.backblaze.com/blog/why-now-is-the-time-for-backb... At these rates, why not use S3? What am I missing?

Amazon S3 is US$0.0275/GB per month at high volumes. I assume Backblaze's operational cost of keeping each Backblaze pod hooked up and online is much much less than US$0.05/GB/month.

Re: Hard Drive Reliability Update – Sep 2014

#73
post #69

Question for the OP here (or for anyone else). Do you burn in new drives before using? I typically will take any new drive and do some type of stress test [1] on it for 18 to 24 hours to see if it fails with that initial constant use. [1] Constant reformatting for example writing 0's to the entire disk 7 times etc.

Yes, these days I'll make sure that it sees a decent amount of power on time and several full drive writes before ever trusting data to it.

Re: Hard Drive Reliability Update – Sep 2014

#74
One of the challenges I have with this analysis is that a 'failure' isn't just that your drive is no longer working, it is that your drive isn't working and you have to go replace it. The operational costs of replacing a drive have three parts, loss of production while the drive is offline, operator time to physically replace the drive and prep it for re-entry into the system, and transactional costs of doing a warranty replacement (filling out the RMA form, getting a valid RMA, shipping the and receiving replacements). We minimize the latter by doing RMAs in batches of 20 but its still a cost across those 20 drives. (and the population of 40 drives which exist as spares are effectively not available for production). It isn't as simple as 'sure drives fail a bit more often but we don't expect to use them that long.'

Re: Hard Drive Reliability Update – Sep 2014

#75

Earlier quoted context omitted.

Would you mind listing some of those time-to-failure and time-to-event methods and how the author might control for them, and which graphs in particular the author should have included?

first, i should've mentioned from the outset that these are really interesting and useful data, and that i'm glad you took the time to generate them. i really wish consumers could systematically report these data in a way that was reliable/trustworthy...! cox proportional hazards models and KM survival curves are the big kahuna with a data set like this. basically, my impression is that you'd want to pretend that you…

Yev from Backblaze here -> We've talked about posting the raw data, but haven't quite decided on that yet. I'll make sure to forward this to Brian so he can take a look and see whether or not any of the above would be feasible for the next one!

Re: Hard Drive Reliability Update – Sep 2014

#76

I love reading these posts from Backblaze, but what I never understand is that they are getting a cost of U$ ~0.05/GB with their storage pods: https://www.backblaze.com/blog/why-now-is-the-time-for-backb... At these rates, why not use S3? What am I missing?

Amazon S3 is US$0.0275/GB per month at high volumes. I assume Backblaze's operational cost of keeping each Backblaze pod hooked up and online is much much less than US$0.05/GB/month.

You're right, sorry. I'm so used to seeing U$/month I didn't realize it was the per GB cost for building the pod.

Re: Hard Drive Reliability Update – Sep 2014

#77
I really want to applaud backblaze for publishing these reports and stats. Too many companies closely guard this information that really helps the larger community. Based on the previous blogs from backblaze, when I built out our new hadoop cluster, I purchased 1450 Hitachi drives. I plan to gather our failure rates and publish them as backblaze does. Thanks for blazing the path!

Re: Hard Drive Reliability Update – Sep 2014

#78
post #9

My main Linux box has quite a few hard drives in it from a large range of time. About 4 weeks ago the oldest of them all died: it is from 2007, so about 7 years old, which I think is pretty good for a consumer drive that's on 24/7. It was a Western Digital Caviar SE WD3200JB, 320GB. I replaced it with a 2TB drive. [No lost data, I do daily backups.]

I wonder what media people use to backup xTB NAS. Tapes ?

Re: Hard Drive Reliability Update – Sep 2014

#79
post #14

Since annual failure rate is a function mostly of age, it would be interesting to see a line chart of cumulative failure rate vs age. But since new drives are continually being added to the population, there would be fewer drives in the data set as you moved up each curve. I guess you could calculate confidence intervals at quarterly intervals, and so the error bars would get larger as age increases and 'n' decreases…

Here's a blog post giving a tutorial on survival analysis using a relatively recent (and IMO very compelling) python library. Could be what you're looking for.

http://camdp.com/blogs/lifelines-survival-analysis-python

Re: Hard Drive Reliability Update – Sep 2014

#80
post #56

How times have changed; Seagate used to be (or at least have the reputation of being) the most reliable and Hitachi the least.

I choose drive brands almost exclusively based on warranty duration. After Seagate bought Maxtor, they started lowering warranties on all of their drives. The results, as you can see, were predictable.

Their "enterprise" disks (except for Terascale HDD/Constellation CS models) apparently still have 5 year warranties†, and I just had to exchange a big Constellation drive, so I can attest their warranty service is still very good.

†They apparently dropped this to 3 years for the first 6 months of 2012 at least for the big drives who's technology base is consumer ... I suspect that change was not well received.

Unfortunately, you are generally correct that we've seen a race to the bottom.

Post reply on HN