Live data from Hacker News

Hard Drive Reliability Update – Sep 2014

backblaze.com

11–20 of 168 posts

Re: Hard Drive Reliability Update – Sep 2014

#11
I manage a computation cluster for an oil and gas exploration company. We have a 50% failure (and rising!) of Seagate Constellation drives in 250GB, 1TB and 2TB configurations. My sample size is fairly small at a few hundred drives but man does it keep me busy.

Re: Hard Drive Reliability Update – Sep 2014

#12

Not that I use even 0.001% of the disks that BackBlaze go through, but my anecdata suggests the same. The only dead hard disks I have on my desk at the moment are Seagate, and they dominate the disks I've sent back in the last few years. However, they are cheap, and they do honour their warranties. Would just be nice if they didn't have to quite so much.

I think my new favorite phrase is going to be "anecdata". Love that.

Re: Hard Drive Reliability Update – Sep 2014

#13
post #9

My main Linux box has quite a few hard drives in it from a large range of time. About 4 weeks ago the oldest of them all died: it is from 2007, so about 7 years old, which I think is pretty good for a consumer drive that's on 24/7. It was a Western Digital Caviar SE WD3200JB, 320GB. I replaced it with a 2TB drive. [No lost data, I do daily backups.]

I just replaced a WD Black 640 3 years into its life. Blacks have a 5 year so they RMA'd it and sent me a 750 in its place.

I do however have a Seagate drive laying around somewhere that has almost 10 years on it and it still functions flawlessly. But it is admittedly smaller given its age and that may contribute to lifespan. Either that or Seagate has slipped in the last decade.

Re: Hard Drive Reliability Update – Sep 2014

#14
Since annual failure rate is a function mostly of age, it would be interesting to see a line chart of cumulative failure rate vs age. But since new drives are continually being added to the population, there would be fewer drives in the data set as you moved up each curve.

I guess you could calculate confidence intervals at quarterly intervals, and so the error bars would get larger as age increases and 'n' decreases.

How would you calculate the CI for failure rate? It's not binomial or poisson, since failure rate goes to 1 over time...

A little searching turns up http://rmod.ee.duke.edu/statistics.htm which I'm sure completely explains how to do this... (rolls eyes). I hate that this is how statistics is commonly taught. Knowing which distribution to use and applying it correctly can actually be intuitive if taught properly. It doesn't always need to be an exercise in alphabet soup / deriving from base principles.

Re: Hard Drive Reliability Update – Sep 2014

#15

Biggest takeaway was at the end, with the "enterprise" drives being slightly less reliable than the consumer ones at half the cost.

And thats whats weird, who is the audience for the article? I know what enterprise drive means and I know the author is smart and knows what it means, so why the weird implications in the article that have nothing to do with "enterprise"?

For those not "in the know" the hardware is the same, but desktop firmware drives will sit there for 10 seconds or whatever it is beating the drive when there's a read (or write) fail on the assumption that if your machine only has one drive you're better off trying as hard as possible to keep retrying until it works, and possibly the slowness will motivate them to replace (god forbid an end user have backups lol)

Enterprise firmware, when it has a soft fail, just croaks as fast as possible. That lets the raid array hurry up and do its thing, or maybe even higher level replication do its thing.

(edited to add the old startup adage of "fail quickly". Thats what enterprise drives do to keep overall array latency low, which is counter productive for consumer non-array drives)

Aside from the firmware load the prices are different because usually enterprise has better guarantee and better service and unlike consumer drives which statistically are never replaced under guarantee so you can claim anything on paper for marketing purposes it won't cost anything, enterprise drives WILL get replaced and there will be a papertrail etc. So the guarantee for an enterprise drive actually costs something.

Sometimes the firmware has some other subtle differences like how it handles recalibrates and scrubs (consumer home drives are like "too bad you get to wait on my schedule" and again, enterprise will go to some effort to eliminate array latency)

My guess is the article is subtle astroturf by the winning drive mfgr?

Re: Hard Drive Reliability Update – Sep 2014

#16

I manage a computation cluster for an oil and gas exploration company. We have a 50% failure (and rising!) of Seagate Constellation drives in 250GB, 1TB and 2TB configurations. My sample size is fairly small at a few hundred drives but man does it keep me busy.

Were they all purchased at the same time? Sounds more like a faulty batch or issue with the environment for a rate that high.

Re: Hard Drive Reliability Update – Sep 2014

#18

All my WD and Seagate drives have failed within two years of use. Call me the luckiest.

My small datacenter results mimic BackBlaze too. Dead/dying seagates all over the place. So I notified management that we will only be purchasing Hitachi drives from now on. I have a BackBlaze server that I recently converted to FreeBSD & ZFS. I love the drive-density that Backblaze offers but I HATE the lack of physical notification when a drive dies.

Most ofther file-servers have a front-facing drive caddy, that usually has LEDs on the front to indicate disk access or errors. This is great because you can walk into the datacenter, and SEE which disk has failed. With the backBlaze system you can get /dev/DriveID but not know where in the array that particular disk is.

Re: Hard Drive Reliability Update – Sep 2014

#19
post #15

Biggest takeaway was at the end, with the "enterprise" drives being slightly less reliable than the consumer ones at half the cost.

And thats whats weird, who is the audience for the article? I know what enterprise drive means and I know the author is smart and knows what it means, so why the weird implications in the article that have nothing to do with "enterprise"? For those not "in the know" the hardware is the same, but desktop firmware drives will sit there for 10 seconds or whatever it is beating the drive when there's a read (or write) fa…

I'm not sure who the audience is, but thank you for your explanation of the differences between consumer and enterprise drive firmwares. That runs counter to my naive assumptions, which was that "enterprise" targeted things tend to be of higher reliability and quality than the consumer-targeted alternatives.

I imagine someone in the drive / enterprise storage business would know better, but many of us might not have known that.

Re: Hard Drive Reliability Update – Sep 2014

#20
post #16

I manage a computation cluster for an oil and gas exploration company. We have a 50% failure (and rising!) of Seagate Constellation drives in 250GB, 1TB and 2TB configurations. My sample size is fairly small at a few hundred drives but man does it keep me busy.

Were they all purchased at the same time? Sounds more like a faulty batch or issue with the environment for a rate that high.

They've been purchased over a span of a few years so I doubt that's the case. They did have a single exposure to 110F ambient temperature for a few hours when the A/C to the server room went out which may be a contributing factor.
Post reply on HN