Live data from Hacker News

Backblaze Hard Drive Stats for 2018

backblaze.com

11–20 of 165 posts

Re: Backblaze Hard Drive Stats for 2018

#11
post #10

Earlier quoted context omitted.

I'm not sure how that would be useful, since terabytes don't fail. When a drive fails, it's effectively a brick with no terabytes.

I was thinking as one failure of a 100TB disk has a very different impact of 10 failures of 1TB disks. It'd give some idea on how much data is lost due failures, no?

Yev from Backblaze here -> Not sure if you'd get that metric from that data. We use Reed-Solomon erasure coding (https://www.backblaze.com/blog/reed-solomon/) to make sure that data is "rebuilt" should we lose drives (which happens all the time).

Re: Backblaze Hard Drive Stats for 2018

#12
post #10

Earlier quoted context omitted.

I'm not sure how that would be useful, since terabytes don't fail. When a drive fails, it's effectively a brick with no terabytes.

I was thinking as one failure of a 100TB disk has a very different impact of 10 failures of 1TB disks. It'd give some idea on how much data is lost due failures, no?

[deleted]

Re: Backblaze Hard Drive Stats for 2018

#13
post #3

HGST look like the best, but they don't have the quantities of the Seagate, makes me wonder if these numbers are skewed :/

They're not going to be skewed. HGST disks are usually more expensive, so that limits the quantity that they buy. The seagate disks are usually cheaper but have a slightly higher (except when it's a brand new line) failure rate. When filling out a single server I go with the HGST disks because the premium price and quality means fewer failures, but it's more cost effective to go for lots of seagate disks when you hav…

I built a server a few years ago, and I determined it was more cost effective for a drive to fail than to use HGST. It would be more inconvenient, but having a drive fail on a home server with only 6 drives didn't seem very likely anyway.

Re: Backblaze Hard Drive Stats for 2018

#14
post #10

Earlier quoted context omitted.

I'm not sure how that would be useful, since terabytes don't fail. When a drive fails, it's effectively a brick with no terabytes.

I was thinking as one failure of a 100TB disk has a very different impact of 10 failures of 1TB disks. It'd give some idea on how much data is lost due failures, no?

I suppose, but there’s no such thing as a single HDD that stores 100TB. The biggest you can get currently are (I believe) 14TB helium-filled drives.

Re: Backblaze Hard Drive Stats for 2018

#15
post #13

Earlier quoted context omitted.

They're not going to be skewed. HGST disks are usually more expensive, so that limits the quantity that they buy. The seagate disks are usually cheaper but have a slightly higher (except when it's a brand new line) failure rate. When filling out a single server I go with the HGST disks because the premium price and quality means fewer failures, but it's more cost effective to go for lots of seagate disks when you hav…

I built a server a few years ago, and I determined it was more cost effective for a drive to fail than to use HGST. It would be more inconvenient, but having a drive fail on a home server with only 6 drives didn't seem very likely anyway.

Particularly if it fails within the warranty period.

Re: Backblaze Hard Drive Stats for 2018

#16
Anything interesting in this one? I've stopped reading them because they all seem to be "Seagates fail kind of a lot but we use them because reasons. HGST doesn't fail a lot, but we also have statistically insignificant numbers of them, so ."

Re: Backblaze Hard Drive Stats for 2018

#17

Anything interesting in this one? I've stopped reading them because they all seem to be "Seagates fail kind of a lot but we use them because reasons. HGST doesn't fail a lot, but we also have statistically insignificant numbers of them, so ."

No, the opportunity to do a better analysis hasn't gone away.

Re: Backblaze Hard Drive Stats for 2018

#18
post #2

Would be interesting to also have metrics on failure per TB storage.

I’m not sure that it’s especially useful to measure that way, which is why they wouldn’t report it. The chance that a given GB of data is on a failed disk is equal to the disk failure rate, regardless of disk size (>1GB).

For large deployments, the concern is between failure rates and the amount of time it takes to rebuild data from a failed disk.

For small deployments, my main concerns are whether disk failure takes a machine or volume out, causing availability problems.

I’m trying to figure where failures per GB would be how you would choose, what scenario we’re you thinking of?

Re: Backblaze Hard Drive Stats for 2018

#20
post #3

HGST look like the best, but they don't have the quantities of the Seagate, makes me wonder if these numbers are skewed :/

My experience, over decades, but now coming up on half that ago, was that IBM/Hitachi/HGST are definitely worth it. I used to run a small hosting business, which had hundreds of discs, and consulting business which had clients with hundreds more discs.

Seagates, in general, could always be expected to fail in a 3 year span. Often multiple times. IBM/Hitachi/HGST, especially the UltraStar enterprise line, would basically never fail. Over 20 years and hundreds of discs concurrently running (switching out as opportunity arose), we probably RMAed single digits of HGST drives.

In comparison, while I had a few pockets of Seagate drives that didn't fail (I had 6 in a storage server at home that never gave me problems), we could generally expect 5% of Seagate drives that we had running to fail in a given year. Enterprise or regular didn't really seem to matter.

For us, replacing a drive was fairly expensive, using a limited resource (our time). So we gravitated to HGST almost exclusively.

But: We also did an extensive burn-in process before a drive went into production. Basically: "badblocks -svw" for a week. He noticed that we had some drives that would fall out of RAID arrays, but if we ran badblocks on would never report an error. My theory was that there were some marginal sectors that would bit-rot. Running a week of badblocks would exercise those and allow badblock remapping to remove them from use.

Remember the IBM "Deathstar" 75GXP? We even had good luck with those. I had one of them start freaking out, and I was aware of the "Deathstar" name, so I went to replace it with another drive I had on hand. When I pulled it out to replace it I realized it was HOT. Not in the grand scheme of things, but definitely hot to the touch. I looked up the temp specs and it was clearly above that. The box it was in had 2 5-inch bays that didn't have the covers on them. I covered those bays, turned the machine back on after re-installing the Deathstar, and the drive continued operating for another 3-5 years with no problem. Made me wonder if improper cooling was the source of those reports.

Post reply on HN