Live data from Hacker News

Hard Drive Reliability Update – Sep 2014

backblaze.com

51–60 of 168 posts

Re: Hard Drive Reliability Update – Sep 2014

#51

How times have changed; Seagate used to be (or at least have the reputation of being) the most reliable and Hitachi the least.

I don't know if I was just lucky or what, but I've always had far worse experiences with Seagates than WDs.

I got into building PCs on the side in about '96 or so and did break/fix support through college - it always seemed (I didn't keep records - I can't quantify) that Seagates accounted for more than their share of the "it won't start and is making a clicking noise" cases. So much so, that I have recommended against Seagates and towards WDs for as long as I can remember (15+yrs).

I've only had a small amount of direct experience with Hitachi drives, so I can't speak to them too much.. Though I do have a Hitachi in an enclosure that's been kicking around since about 2004, so I guess that says something...

Re: Hard Drive Reliability Update – Sep 2014

#52

How times have changed; Seagate used to be (or at least have the reputation of being) the most reliable and Hitachi the least.

At some point, Hitachi Diskstars were referred to as "Death Stars", and that was all I knew about disk reliability. It is great to have some real information. Dell, Google, and Amazon can never write reports like this because the vendor relationship is important. Because these guys have no relationship and are buying consumer disks, the world finally gets brand level reliability reports. Kudos to Backblaze.

I think that was one IBM model way before Hitachi acquired that division from Big Blue; they've otherwise been pretty good.

Re: Hard Drive Reliability Update – Sep 2014

#53

There are well established methods for time-to-failure and time-to-event data not used here. The author makes no effort to control for the multiple, obvious biases created by the analytical approach employed. A few simple graphs would give a much more telling view of these data.

Would you mind listing some of those time-to-failure and time-to-event methods and how the author might control for them, and which graphs in particular the author should have included?

first, i should've mentioned from the outset that these are really interesting and useful data, and that i'm glad you took the time to generate them. i really wish consumers could systematically report these data in a way that was reliable/trustworthy...!

cox proportional hazards models and KM survival curves are the big kahuna with a data set like this. basically, my impression is that you'd want to pretend that you're doing a cohort study and essentially analyze it like a big clinical trial.

and re: graphing, since you have sub-samples of the different drive models and relatively different but still small numbers of models for each brand, the big take-away graph comparing manufacturers leaves out a lot of useful nuances that could be pulled out of the table, e.g. that most of the hitachi drives are newer and you have fewer of them. it also doesn't portray how consistent failure rates are across models produced by the same brand might be, and it looks like there's a significant range of failure rates across different seagate models. so even just seeing IQRs of pooled failure rates for each manufacturer in box plots might be eye opening..

immortal time bias may be something to consider here as well...when you have some sample groups that mostly include newer individual drives that you have not yet had for a year on average, subtle differences in how the failure event is described can make a big difference in the conclusions you draw...especially in terms of uncertainty. if i have 10,000 hitachi drives and 5 of them fail in the first 6 months, the robustness of conclusions i can draw from those data are different in some important ways from similar insights drawn from a sample of 1,000 hitachi drives i've used for 5 years.

it's also not clear to me how you've dealt with replacement drives. based on what i gather from the post (and i could be totally wrong), if you have a bunch of one model of a drive failing and then get replacements for them... some might argue that refurbed drives should be analyzed almost like a separate drive model, since they often have physical differences compared to those bought via retail channels.

i'd be quite curious to dig into these data a bit further if you're willing to post the original data set...

thanks again for posting information that's quite useful and interesting

Re: Hard Drive Reliability Update – Sep 2014

#54

How times have changed; Seagate used to be (or at least have the reputation of being) the most reliable and Hitachi the least.

Related: Intel's SSD is the only SSD I've ever lost data with. It happened due to a power outage that bricked the drive, and the only reason it happened was because of a flaw in their firmware, not their drive, which they hushed up. I was very surprised because Intel had the reputation of being the best SSD at the time. (It was the 300-something series.)

My Kingston SSD is already paying its aging tax.

Re: Hard Drive Reliability Update – Sep 2014

#55
post #15

Biggest takeaway was at the end, with the "enterprise" drives being slightly less reliable than the consumer ones at half the cost.

And thats whats weird, who is the audience for the article? I know what enterprise drive means and I know the author is smart and knows what it means, so why the weird implications in the article that have nothing to do with "enterprise"? For those not "in the know" the hardware is the same, but desktop firmware drives will sit there for 10 seconds or whatever it is beating the drive when there's a read (or write) fa…

Enterprise firmware, when it has a soft fail, just croaks as fast as possible. That lets the raid array hurry up and do its thing, or maybe even higher level replication do its thing.

Except that also only works "most of the time".

Anyone who works at scale with drives and (RAID-)Controllers knows that even "enterprise" drives can and do take entire controllers down.

It's not uncommon to lose a full set of daisy chained JBODs to a single disk acting funny.

Consequently, and since storage clusters have to be redundant at the node-level anyway, it makes a lot of sense to skip the markup for enterprise firmwares and instead design for quick node failure detection and ejection (short timeouts).

Re: Hard Drive Reliability Update – Sep 2014

#56

How times have changed; Seagate used to be (or at least have the reputation of being) the most reliable and Hitachi the least.

I choose drive brands almost exclusively based on warranty duration. After Seagate bought Maxtor, they started lowering warranties on all of their drives. The results, as you can see, were predictable.

Re: Hard Drive Reliability Update – Sep 2014

#57
post #10

I wish there was something similar for SSDs.

After a few hundred drives, our anecdata is that failure over time on SSDs is generally related to drive endurance. Make sure you use a SMART utility which can read (and translate to English) the current net usage of the drive. Throw them away when you get to 100% usage.

I recently examined a set of Crucial m4s which were at 130% of usage. There was no lost data, but write bandwidth was hilariously bad (around 10-20MB/s).

Re: Hard Drive Reliability Update – Sep 2014

#58
post #14

Since annual failure rate is a function mostly of age, it would be interesting to see a line chart of cumulative failure rate vs age. But since new drives are continually being added to the population, there would be fewer drives in the data set as you moved up each curve. I guess you could calculate confidence intervals at quarterly intervals, and so the error bars would get larger as age increases and 'n' decreases…

Couldn't a Weibull Distribution be used to determine the shape parameter (how the failure rate changes with time)?

EDIT: More fine grained data is probably needed, granted.

Re: Hard Drive Reliability Update – Sep 2014

#59

How times have changed; Seagate used to be (or at least have the reputation of being) the most reliable and Hitachi the least.

I used to love Seagate. Then I got bit by the 7200.11 firmware problems, which was one of the worst hard drive experiences I've ever had. Never again.

Re: Hard Drive Reliability Update – Sep 2014

#60
post #15

Earlier quoted context omitted.

And thats whats weird, who is the audience for the article? I know what enterprise drive means and I know the author is smart and knows what it means, so why the weird implications in the article that have nothing to do with "enterprise"? For those not "in the know" the hardware is the same, but desktop firmware drives will sit there for 10 seconds or whatever it is beating the drive when there's a read (or write) fa…

unlike consumer drives which statistically are never replaced under guarantee They're never replaced because regular consumers statistically don't ask for replacement and simply eat the loss?

i am a statistical anomaly, then, as someone who tracks warranty expirations in a spreadsheet and always makes sure to wipe & sell old drives before they go out of warranty. in the 20 years i've owned hard drives, i've always had failed drives replaced!

drives with longer warranties generally cost more in the consumer segment, and data like this could help people identify the "sweet spot" balancing cost with risk of failure, warranty length, and depreciation curves for storage. my "ancedata" suggests somewhere between 2 and 3 year warranties seems about right for consumer drives these days...

Post reply on HN