Live data from Hacker News

Hard Drive Reliability Update – Sep 2014

backblaze.com

81–90 of 168 posts

Re: Hard Drive Reliability Update – Sep 2014

#81

I manage a computation cluster for an oil and gas exploration company. We have a 50% failure (and rising!) of Seagate Constellation drives in 250GB, 1TB and 2TB configurations. My sample size is fairly small at a few hundred drives but man does it keep me busy.

50% annual failure rate, or 50% total, over several years?

50% total over 2-5 years.

Re: Hard Drive Reliability Update – Sep 2014

#82
post #66

Earlier quoted context omitted.

unlike consumer drives which statistically are never replaced under guarantee They're never replaced because regular consumers statistically don't ask for replacement and simply eat the loss?

"They're never replaced because regular consumers statistically don't ask for replacement and simply eat the loss?" There is another paradox as well. Some people won't ask for replacement because assuming you have to send in the bricked drive there is the chance that someone might get at your data somehow. What about that? (It's why I would never send in a drive that has failed.) [1] [1] My assumption is that I would…

I encrypt my drives anyway. If they can reanimate the drive and crack the encryption, well, they deserve to see the data.

Re: Hard Drive Reliability Update – Sep 2014

#83
post #15

Earlier quoted context omitted.

And thats whats weird, who is the audience for the article? I know what enterprise drive means and I know the author is smart and knows what it means, so why the weird implications in the article that have nothing to do with "enterprise"? For those not "in the know" the hardware is the same, but desktop firmware drives will sit there for 10 seconds or whatever it is beating the drive when there's a read (or write) fa…

Are you claiming: o Enterprise and Consumer Drives are the Same Hardware. o Enterprise Firmware causes the drive to fail fast on physical errors rather than endlessly retrying. o Enterprise Drives are more likely to be RMA'd? Do you have any citations, evidence, reports, articles, white papers, research - anything (beyond random anecdotes, or miscellaneous blog entries) to back up these claims?

>Enterprise Firmware causes the drive to fail fast on physical errors rather than endlessly retrying. TLER

http://features.techworld.com/storage/1019/what-is-time-limi...

TLER is exactly what you want in a RAID. You have another copy of the data in question on the array, why fight for it on a disk when the data may be corrupt anyway.

Re: Hard Drive Reliability Update – Sep 2014

#84

How times have changed; Seagate used to be (or at least have the reputation of being) the most reliable and Hitachi the least.

At some point, Hitachi Diskstars were referred to as "Death Stars", and that was all I knew about disk reliability. It is great to have some real information. Dell, Google, and Amazon can never write reports like this because the vendor relationship is important. Because these guys have no relationship and are buying consumer disks, the world finally gets brand level reliability reports. Kudos to Backblaze.

If I remember correctly they didn't name names, but Google did publish a very useful paper on their experiences with disks. The two big takeaways were that in only half the disks the self testing signaled a future failure, and manufactures seemed to have solved heat problems.

Re: Hard Drive Reliability Update – Sep 2014

#85
post #16

Earlier quoted context omitted.

Were they all purchased at the same time? Sounds more like a faulty batch or issue with the environment for a rate that high.

They've been purchased over a span of a few years so I doubt that's the case. They did have a single exposure to 110F ambient temperature for a few hours when the A/C to the server room went out which may be a contributing factor.

I have found that to be very hard on disks. Lost a AC unit in summer and over the next month had a much higher rate of disk loss.

Re: Hard Drive Reliability Update – Sep 2014

#86

How times have changed; Seagate used to be (or at least have the reputation of being) the most reliable and Hitachi the least.

I used to love Seagate. Then I got bit by the 7200.11 firmware problems, which was one of the worst hard drive experiences I've ever had. Never again.

Unfortunately, there are so few choice now. :-(

Re: Hard Drive Reliability Update – Sep 2014

#87
post #83

Earlier quoted context omitted.

Are you claiming: o Enterprise and Consumer Drives are the Same Hardware. o Enterprise Firmware causes the drive to fail fast on physical errors rather than endlessly retrying. o Enterprise Drives are more likely to be RMA'd? Do you have any citations, evidence, reports, articles, white papers, research - anything (beyond random anecdotes, or miscellaneous blog entries) to back up these claims?

>Enterprise Firmware causes the drive to fail fast on physical errors rather than endlessly retrying. TLER http://features.techworld.com/storage/1019/what-is-time-limi... TLER is exactly what you want in a RAID. You have another copy of the data in question on the array, why fight for it on a disk when the data may be corrupt anyway.

Fair enough - that's a good article, though it explicitly calls out "RE" drives - "RAID Edition" - which I can imagine having particularly properties associated with "RAID" behavior.

What I'm interesting in hearing, (honestly - I not doubting right now, just interested in being educated) - is if anyone authoritative has described "Enterprise" drives as having these behaviors.

Re: Hard Drive Reliability Update – Sep 2014

#88
post #9

My main Linux box has quite a few hard drives in it from a large range of time. About 4 weeks ago the oldest of them all died: it is from 2007, so about 7 years old, which I think is pretty good for a consumer drive that's on 24/7. It was a Western Digital Caviar SE WD3200JB, 320GB. I replaced it with a 2TB drive. [No lost data, I do daily backups.]

I wonder what media people use to backup xTB NAS. Tapes ?

I back up my RAIDs to JBODs - so hard drive backups for hard drives, just fewer. I only use tape for vital and frequently backed up things - not entire drives, and generally not large.

Re: Hard Drive Reliability Update – Sep 2014

#90

One of the challenges I have with this analysis is that a 'failure' isn't just that your drive is no longer working, it is that your drive isn't working and you have to go replace it. The operational costs of replacing a drive have three parts, loss of production while the drive is offline, operator time to physically replace the drive and prep it for re-entry into the system, and transactional costs of doing a warra…

That's a good point but also requires a more complicated analysis since it requires you to account for your ops capacity (e.g. the marginal cost increase of doing 10 drives instead of 5 is probably a lot less than 2 unless your ops team is almost completely booked) and also requires some guessing about whether there actually is a better option. They addressed that second point with the discussion of enterprise drives, which certainly matched my own experience.
Post reply on HN