Live data from Hacker News

What SMART Stats Tell Us About Hard Drives

backblaze.com

1–10 of 71 posts

Re: What SMART Stats Tell Us About Hard Drives

#2
SMART is a fantastic exercise in sensitivity and specificity. As backblaze is showing with this data, SMART stats have poor sensitivity, but what's much worse for those who run big fleets of drives is their poor specificity. Lots of healthy drives are reported unhealthy by SMART. If I'm running a gold-plated database server, that doesn't matter. A couple of extra planned drive replacements is a small price to pay for avoiding unplanned failures. If I'm running a huge drive cluster, it's much, much more expensive.

Take Backblaze's 0.01% for a group of four failures. That's replacing an extra 100 drives per million, at random, and only getting the benefit of correctly predicting failures 10.4% of the time.

This is great data to have.

Re: What SMART Stats Tell Us About Hard Drives

#3
post #2

SMART is a fantastic exercise in sensitivity and specificity. As backblaze is showing with this data, SMART stats have poor sensitivity, but what's much worse for those who run big fleets of drives is their poor specificity. Lots of healthy drives are reported unhealthy by SMART. If I'm running a gold-plated database server, that doesn't matter. A couple of extra planned drive replacements is a small price to pay for…

Nice analysis.

Re: What SMART Stats Tell Us About Hard Drives

#6
Interesting, thanks for posting! Could you talk quickly about why it's interesting to predict drive failure? Is it to understand how many replacement drives you might need to order in the short term, or is there value beyond stock management of drives?

Re: What SMART Stats Tell Us About Hard Drives

#8
post #2

SMART is a fantastic exercise in sensitivity and specificity. As backblaze is showing with this data, SMART stats have poor sensitivity, but what's much worse for those who run big fleets of drives is their poor specificity. Lots of healthy drives are reported unhealthy by SMART. If I'm running a gold-plated database server, that doesn't matter. A couple of extra planned drive replacements is a small price to pay for…

Thresholds are often useful with these kinds of stats. Aka a drive moving one sector might mean nothing, but moving 30 in a week could be great predictor. Further they only have 70k drives across a range of product lines so what predicts drive X failing very well might say little about drive Y.

PS: Rememebr all RAM gets bit flit errors over time. Which is one of the reasons rebooting is often so useful, but also means one off errors are often meaningless.

Re: What SMART Stats Tell Us About Hard Drives

#9
post #6

Interesting, thanks for posting! Could you talk quickly about why it's interesting to predict drive failure? Is it to understand how many replacement drives you might need to order in the short term, or is there value beyond stock management of drives?

Not OP, but:

Perfect predictability would obviously be beneficial, in that you could get by without any redundancy. But even imperfect predicability can help you reduce the required number of drives for a set level of security.

Re: What SMART Stats Tell Us About Hard Drives

#10
post #2

SMART is a fantastic exercise in sensitivity and specificity. As backblaze is showing with this data, SMART stats have poor sensitivity, but what's much worse for those who run big fleets of drives is their poor specificity. Lots of healthy drives are reported unhealthy by SMART. If I'm running a gold-plated database server, that doesn't matter. A couple of extra planned drive replacements is a small price to pay for…

[deleted]
Post reply on HN