Live data from Hacker News

Hard Drive Reliability Update – Sep 2014

backblaze.com

31–40 of 168 posts

Re: Hard Drive Reliability Update – Sep 2014

#31
post #26
post #10

I wish there was something similar for SSDs.

There is, roughly. This is a lab test as opposed to production monitoring, but it is still interesting. Obviously SSDs and HDDs are quite different technically, so the appropriate test is quite different. They write data continuously to see how long they last. We've now written over a petabyte, and only half of the SSDs remain. Three drives failed at different points—and in different ways—before reaching the 1PB mile…

Latest results: http://techreport.com/review/27062/the-ssd-endurance-experim...

Re: Hard Drive Reliability Update – Sep 2014

#33

How times have changed; Seagate used to be (or at least have the reputation of being) the most reliable and Hitachi the least.

At some point, Hitachi Diskstars were referred to as "Death Stars", and that was all I knew about disk reliability. It is great to have some real information.

Dell, Google, and Amazon can never write reports like this because the vendor relationship is important. Because these guys have no relationship and are buying consumer disks, the world finally gets brand level reliability reports. Kudos to Backblaze.

Re: Hard Drive Reliability Update – Sep 2014

#35
There are well established methods for time-to-failure and time-to-event data not used here. The author makes no effort to control for the multiple, obvious biases created by the analytical approach employed. A few simple graphs would give a much more telling view of these data.

Re: Hard Drive Reliability Update – Sep 2014

#36
post #15

Earlier quoted context omitted.

And thats whats weird, who is the audience for the article? I know what enterprise drive means and I know the author is smart and knows what it means, so why the weird implications in the article that have nothing to do with "enterprise"? For those not "in the know" the hardware is the same, but desktop firmware drives will sit there for 10 seconds or whatever it is beating the drive when there's a read (or write) fa…

unlike consumer drives which statistically are never replaced under guarantee They're never replaced because regular consumers statistically don't ask for replacement and simply eat the loss?

Yes. I have an old 250 GB sitting on my desk at home. Waste the time arguing on the phone and getting a RMA and package it up and drive out to ship it back to get a "new" 250G drive (uh ... thanks?), or order a new 2TB, about the same cost. And this is only with technologically inclined consumers, most are just going to return the whole computer or get a new computer.

At a business that is probably some MBA metric and has been budgeted for and is salaried anyway, and you need to prove I'm not just taking drives home to put in my basement server, so they have to be destroyed or sent back, and there may or may not be PCI / CPNI type concerns, so yeah, they get sent back.

Re: Hard Drive Reliability Update – Sep 2014

#37
post #9

My main Linux box has quite a few hard drives in it from a large range of time. About 4 weeks ago the oldest of them all died: it is from 2007, so about 7 years old, which I think is pretty good for a consumer drive that's on 24/7. It was a Western Digital Caviar SE WD3200JB, 320GB. I replaced it with a 2TB drive. [No lost data, I do daily backups.]

I just replaced a WD Black 640 3 years into its life. Blacks have a 5 year so they RMA'd it and sent me a 750 in its place. I do however have a Seagate drive laying around somewhere that has almost 10 years on it and it still functions flawlessly. But it is admittedly smaller given its age and that may contribute to lifespan. Either that or Seagate has slipped in the last decade.

Seagate apparently has slipped in the last 2-3 years. (Could this be related to the hard disk shortage due to flooding in Malaysia?)

Hard drive failures tend to happen more for drives that are power cycled a lot, and for drives that undergo big swings in temperature (even when the temps are all within the rated temp range).

Re: Hard Drive Reliability Update – Sep 2014

#38
post #15

Earlier quoted context omitted.

And thats whats weird, who is the audience for the article? I know what enterprise drive means and I know the author is smart and knows what it means, so why the weird implications in the article that have nothing to do with "enterprise"? For those not "in the know" the hardware is the same, but desktop firmware drives will sit there for 10 seconds or whatever it is beating the drive when there's a read (or write) fa…

unlike consumer drives which statistically are never replaced under guarantee They're never replaced because regular consumers statistically don't ask for replacement and simply eat the loss?

With a shorter guarantee, consumer drives have less time in which to fail and be replaced without charge.

Re: Hard Drive Reliability Update – Sep 2014

#39

I manage a computation cluster for an oil and gas exploration company. We have a 50% failure (and rising!) of Seagate Constellation drives in 250GB, 1TB and 2TB configurations. My sample size is fairly small at a few hundred drives but man does it keep me busy.

50% annual failure rate, or 50% total, over several years?

Re: Hard Drive Reliability Update – Sep 2014

#40

There are well established methods for time-to-failure and time-to-event data not used here. The author makes no effort to control for the multiple, obvious biases created by the analytical approach employed. A few simple graphs would give a much more telling view of these data.

Would you mind listing some of those time-to-failure and time-to-event methods and how the author might control for them, and which graphs in particular the author should have included?
Post reply on HN