Hard Drive Reliability Update – Sep 2014
11–20 of 168 posts
Re: Hard Drive Reliability Update – Sep 2014
#12Not that I use even 0.001% of the disks that BackBlaze go through, but my anecdata suggests the same. The only dead hard disks I have on my desk at the moment are Seagate, and they dominate the disks I've sent back in the last few years. However, they are cheap, and they do honour their warranties. Would just be nice if they didn't have to quite so much.
Re: Hard Drive Reliability Update – Sep 2014
#13My main Linux box has quite a few hard drives in it from a large range of time. About 4 weeks ago the oldest of them all died: it is from 2007, so about 7 years old, which I think is pretty good for a consumer drive that's on 24/7. It was a Western Digital Caviar SE WD3200JB, 320GB. I replaced it with a 2TB drive. [No lost data, I do daily backups.]
I do however have a Seagate drive laying around somewhere that has almost 10 years on it and it still functions flawlessly. But it is admittedly smaller given its age and that may contribute to lifespan. Either that or Seagate has slipped in the last decade.
Re: Hard Drive Reliability Update – Sep 2014
#14I guess you could calculate confidence intervals at quarterly intervals, and so the error bars would get larger as age increases and 'n' decreases.
How would you calculate the CI for failure rate? It's not binomial or poisson, since failure rate goes to 1 over time...
A little searching turns up http://rmod.ee.duke.edu/statistics.htm which I'm sure completely explains how to do this... (rolls eyes). I hate that this is how statistics is commonly taught. Knowing which distribution to use and applying it correctly can actually be intuitive if taught properly. It doesn't always need to be an exercise in alphabet soup / deriving from base principles.
Re: Hard Drive Reliability Update – Sep 2014
#15Biggest takeaway was at the end, with the "enterprise" drives being slightly less reliable than the consumer ones at half the cost.
For those not "in the know" the hardware is the same, but desktop firmware drives will sit there for 10 seconds or whatever it is beating the drive when there's a read (or write) fail on the assumption that if your machine only has one drive you're better off trying as hard as possible to keep retrying until it works, and possibly the slowness will motivate them to replace (god forbid an end user have backups lol)
Enterprise firmware, when it has a soft fail, just croaks as fast as possible. That lets the raid array hurry up and do its thing, or maybe even higher level replication do its thing.
(edited to add the old startup adage of "fail quickly". Thats what enterprise drives do to keep overall array latency low, which is counter productive for consumer non-array drives)
Aside from the firmware load the prices are different because usually enterprise has better guarantee and better service and unlike consumer drives which statistically are never replaced under guarantee so you can claim anything on paper for marketing purposes it won't cost anything, enterprise drives WILL get replaced and there will be a papertrail etc. So the guarantee for an enterprise drive actually costs something.
Sometimes the firmware has some other subtle differences like how it handles recalibrates and scrubs (consumer home drives are like "too bad you get to wait on my schedule" and again, enterprise will go to some effort to eliminate array latency)
My guess is the article is subtle astroturf by the winning drive mfgr?
Re: Hard Drive Reliability Update – Sep 2014
#16I manage a computation cluster for an oil and gas exploration company. We have a 50% failure (and rising!) of Seagate Constellation drives in 250GB, 1TB and 2TB configurations. My sample size is fairly small at a few hundred drives but man does it keep me busy.
Re: Hard Drive Reliability Update – Sep 2014
#17Re: Hard Drive Reliability Update – Sep 2014
#18All my WD and Seagate drives have failed within two years of use. Call me the luckiest.
Most ofther file-servers have a front-facing drive caddy, that usually has LEDs on the front to indicate disk access or errors. This is great because you can walk into the datacenter, and SEE which disk has failed. With the backBlaze system you can get /dev/DriveID but not know where in the array that particular disk is.
Re: Hard Drive Reliability Update – Sep 2014
#19Biggest takeaway was at the end, with the "enterprise" drives being slightly less reliable than the consumer ones at half the cost.
And thats whats weird, who is the audience for the article? I know what enterprise drive means and I know the author is smart and knows what it means, so why the weird implications in the article that have nothing to do with "enterprise"? For those not "in the know" the hardware is the same, but desktop firmware drives will sit there for 10 seconds or whatever it is beating the drive when there's a read (or write) fa…
I imagine someone in the drive / enterprise storage business would know better, but many of us might not have known that.
Re: Hard Drive Reliability Update – Sep 2014
#20I manage a computation cluster for an oil and gas exploration company. We have a 50% failure (and rising!) of Seagate Constellation drives in 250GB, 1TB and 2TB configurations. My sample size is fairly small at a few hundred drives but man does it keep me busy.
Were they all purchased at the same time? Sounds more like a faulty batch or issue with the environment for a rate that high.