Live data from Hacker News

Hard Drive Reliability Update – Sep 2014

backblaze.com

61–70 of 168 posts

Re: Hard Drive Reliability Update – Sep 2014

#61
post #41

How times have changed; Seagate used to be (or at least have the reputation of being) the most reliable and Hitachi the least.

I'm so ashamed right now. I've been recommending Seagates to everybody who asked for years without updating my fundaments...

Yev from Backblaze here -> No reason to be ashamed! Truth is we still buy Seagates! Most likely they will work well enough in a home environment. All drives fail eventually. If you can get 3-4 years out of one, that's great!

Re: Hard Drive Reliability Update – Sep 2014

#62
I recently bought 3TB Western Digital Red, following their advice from [1], but now I see that it has yearly failure rate of 8.8%, bummer.

Off-topic, but It's a shame that BackBlaze isn't available in some countries, I'd love to use it. What would be the best alternative to it, Tarsnap?

[1] https://www.backblaze.com/blog/what-hard-drive-should-i-buy/

Re: Hard Drive Reliability Update – Sep 2014

#63
post #56

How times have changed; Seagate used to be (or at least have the reputation of being) the most reliable and Hitachi the least.

I choose drive brands almost exclusively based on warranty duration. After Seagate bought Maxtor, they started lowering warranties on all of their drives. The results, as you can see, were predictable.

regressing warranty length against average time to failure in this data set would be interesting...!

Re: Hard Drive Reliability Update – Sep 2014

#64
Used in a small file server, my net failure rate on Seagate's consumer 3TB drives has been over 50% thus far. The pair of their SAS drives I currently have in use have been fine, although both of them are still below a year of service life... Edit: Just checked my drive status, and yet another one has dropped. If I'm doing my math correctly, that's 75% of the drives that weren't DOA...

Re: Hard Drive Reliability Update – Sep 2014

#65
post #14

Since annual failure rate is a function mostly of age, it would be interesting to see a line chart of cumulative failure rate vs age. But since new drives are continually being added to the population, there would be fewer drives in the data set as you moved up each curve. I guess you could calculate confidence intervals at quarterly intervals, and so the error bars would get larger as age increases and 'n' decreases…

Undergrad version: the lifespan of a single drive should be exponentially distributed with some rate lambda. The time to first failure from among k drives would then have an exponential distribution with rate k * lambda. The confidence interval for lambda is: http://en.wikipedia.org/wiki/Exponential_distribution#Confid...

Grad-level version: drive lifetime is exponential with an inhomogeneous (increasing, presumably) rate function lambda(t). Inferring lambda(t) is difficult without additional assumptions on the functional form. But potentially do-able.

Real-life version: None of the classical distribution fit that well. (https://www.usenix.org/legacy/event/fast07/tech/schroeder/sc...)

(This is survival analysis: http://en.wikipedia.org/wiki/Survival_analysis)

Re: Hard Drive Reliability Update – Sep 2014

#66
post #15

Earlier quoted context omitted.

And thats whats weird, who is the audience for the article? I know what enterprise drive means and I know the author is smart and knows what it means, so why the weird implications in the article that have nothing to do with "enterprise"? For those not "in the know" the hardware is the same, but desktop firmware drives will sit there for 10 seconds or whatever it is beating the drive when there's a read (or write) fa…

unlike consumer drives which statistically are never replaced under guarantee They're never replaced because regular consumers statistically don't ask for replacement and simply eat the loss?

"They're never replaced because regular consumers statistically don't ask for replacement and simply eat the loss?"

There is another paradox as well. Some people won't ask for replacement because assuming you have to send in the bricked drive there is the chance that someone might get at your data somehow.

What about that? (It's why I would never send in a drive that has failed.) [1]

[1] My assumption is that I would have to send in the bad drive (and there is no way to reformat or for me to easily destroy what might be on there). Anyone have experience with what happens here?

Re: Hard Drive Reliability Update – Sep 2014

#67
post #14

Since annual failure rate is a function mostly of age, it would be interesting to see a line chart of cumulative failure rate vs age. But since new drives are continually being added to the population, there would be fewer drives in the data set as you moved up each curve. I guess you could calculate confidence intervals at quarterly intervals, and so the error bars would get larger as age increases and 'n' decreases…

It's been 20 years since I took a reliability engineering class but I believe the go to curve is the negative exponential. Here is a link from a quick Googling: http://www.quanterion.com/FAQ/Exponential_Dist.htm

The exponential lifetime distribution is the model for ideal memoryless failures: the probability of failure in a dt interval is independent of current lifetime and have value lambda*dt. Those are as "random" failures as they get. I suppose hard drives are better modelled by a variable failure rate lamba(t), which should have a peak for the first few hours/days, settle down and then start growing quickly after a few months.

Re: Hard Drive Reliability Update – Sep 2014

#69
Question for the OP here (or for anyone else).

Do you burn in new drives before using? I typically will take any new drive and do some type of stress test [1] on it for 18 to 24 hours to see if it fails with that initial constant use.

[1] Constant reformatting for example writing 0's to the entire disk 7 times etc.

Re: Hard Drive Reliability Update – Sep 2014

#70

I love reading these posts from Backblaze, but what I never understand is that they are getting a cost of U$ ~0.05/GB with their storage pods: https://www.backblaze.com/blog/why-now-is-the-time-for-backb... At these rates, why not use S3? What am I missing?

Perhaps being in control of your own destiny?
Post reply on HN