Live data from Hacker News

Hard Drive Reliability Update – Sep 2014

backblaze.com

121–130 of 168 posts

Re: Hard Drive Reliability Update – Sep 2014

#122

How times have changed; Seagate used to be (or at least have the reputation of being) the most reliable and Hitachi the least.

At some point, Hitachi Diskstars were referred to as "Death Stars", and that was all I knew about disk reliability. It is great to have some real information. Dell, Google, and Amazon can never write reports like this because the vendor relationship is important. Because these guys have no relationship and are buying consumer disks, the world finally gets brand level reliability reports. Kudos to Backblaze.

I had four of the three Death Stars I bought fail (yes one of the replacements failed too). 133% failure rate is quite an achievement. They were great drives when they worked.

Re: Hard Drive Reliability Update – Sep 2014

#123
post #15

Biggest takeaway was at the end, with the "enterprise" drives being slightly less reliable than the consumer ones at half the cost.

And thats whats weird, who is the audience for the article? I know what enterprise drive means and I know the author is smart and knows what it means, so why the weird implications in the article that have nothing to do with "enterprise"? For those not "in the know" the hardware is the same, but desktop firmware drives will sit there for 10 seconds or whatever it is beating the drive when there's a read (or write) fa…

Seriously, no. Wrong in almost every way.

Disclaimer: HGST employee.

I'll leave this non-HGST (Intel) link (PDF) for reference:

http://download.intel.com/support/motherboards/server/sb/ent...

Re: Hard Drive Reliability Update – Sep 2014

#124
post #54

Earlier quoted context omitted.

Related: Intel's SSD is the only SSD I've ever lost data with. It happened due to a power outage that bricked the drive, and the only reason it happened was because of a flaw in their firmware, not their drive, which they hushed up. I was very surprised because Intel had the reputation of being the best SSD at the time. (It was the 300-something series.)

My Kingston SSD is already paying its aging tax.

"paying its aging tax" ?

Re: Hard Drive Reliability Update – Sep 2014

#125
post #99
post #77

I really want to applaud backblaze for publishing these reports and stats. Too many companies closely guard this information that really helps the larger community. Based on the previous blogs from backblaze, when I built out our new hadoop cluster, I purchased 1450 Hitachi drives. I plan to gather our failure rates and publish them as backblaze does. Thanks for blazing the path!

Yev from Backblaze -> Thanks! That really is one of our goals with these updates, is for others to join us and start sharing this data. It makes a lot more sense if everyone is doing it, then we can start comparing environments, and all sorts of fun stuff!

For the Seagate 7200.14, have you checked to see if you have the latest firmware on all of them? http://knowledge.seagate.com/articles/en_US/FAQ/223651en

Same for the 7200.11: http://knowledge.seagate.com/articles/en_US/FAQ/207951en

Re: Hard Drive Reliability Update – Sep 2014

#126

How times have changed; Seagate used to be (or at least have the reputation of being) the most reliable and Hitachi the least.

I used to love Seagate. Then I got bit by the 7200.11 firmware problems, which was one of the worst hard drive experiences I've ever had. Never again.

The 7200.11 debacle! It was unbelievable that it happened- as far as I remember, Seagate was the leader in reliability before then, I had created five or six RAIDs with Seagate drives with 100% reliability.

But when they released the 7200.11 versions....ALL 7200.11 models would spin down after idle for X minutes and then start clicking...the data was still intact, but for the drive to work again you had to pull the power.

I unfortunately built an eight drive RAID5 with these drives before the issue was know, which made it very difficult for me to diagnose the issue. (All drives seemed perfectly fine when powered and working for the first 5-10 minutes or whenever they were in use, but as soon as a single drive of the RAID idled, the RAID5 acted like a drive was bad).

I still have a box with eight 1TB Seagate 7200.11 drives that I never updated. I have never bought another Seagate drive for RAID since, always WD or Hitachi.

Re: Hard Drive Reliability Update – Sep 2014

#127
post #77

I really want to applaud backblaze for publishing these reports and stats. Too many companies closely guard this information that really helps the larger community. Based on the previous blogs from backblaze, when I built out our new hadoop cluster, I purchased 1450 Hitachi drives. I plan to gather our failure rates and publish them as backblaze does. Thanks for blazing the path!

Damn how much did that set you back? How much space? SSDs or HDDs? Very curious. What kind of datasets are you analyzing with that much space?

Re: Hard Drive Reliability Update – Sep 2014

#128
post #14

Since annual failure rate is a function mostly of age, it would be interesting to see a line chart of cumulative failure rate vs age. But since new drives are continually being added to the population, there would be fewer drives in the data set as you moved up each curve. I guess you could calculate confidence intervals at quarterly intervals, and so the error bars would get larger as age increases and 'n' decreases…

You can see cumulative failure rate vs age here (last graph)

https://www.backblaze.com/blog/what-hard-drive-should-i-buy/

Re: Hard Drive Reliability Update – Sep 2014

#129
post #99

Earlier quoted context omitted.

Yev from Backblaze -> Thanks! That really is one of our goals with these updates, is for others to join us and start sharing this data. It makes a lot more sense if everyone is doing it, then we can start comparing environments, and all sorts of fun stuff!

For the Seagate 7200.14, have you checked to see if you have the latest firmware on all of them? http://knowledge.seagate.com/articles/en_US/FAQ/223651en Same for the 7200.11: http://knowledge.seagate.com/articles/en_US/FAQ/207951en

Yup, we have top-men that are in charge of keeping an eye on firmware, if we see a need in an update, we'll do it.

Re: Hard Drive Reliability Update – Sep 2014

#130
post #14

Since annual failure rate is a function mostly of age, it would be interesting to see a line chart of cumulative failure rate vs age. But since new drives are continually being added to the population, there would be fewer drives in the data set as you moved up each curve. I guess you could calculate confidence intervals at quarterly intervals, and so the error bars would get larger as age increases and 'n' decreases…

The financial analogy is to think of failure rates in drives as equivalent to defaults on a pool of securities (say mortgages). obviously there are a number of ways to do this but the most simple is that since origination standards for loans vary through time (think build quality and design changes in the drive context) it seems to make sense to think about this from a vintage standpoint. i.e. are seagate drives installed in Dec'12 failing at the same rate as ones installed in March '13 after six months...
Post reply on HN