Live data from Hacker News

SSD will fail at 40k power-on hours (2021)

cisco.com

161–170 of 276 posts

Re: SSD will fail at 40k power-on hours (2021)

#161
post #139

Earlier quoted context omitted.

This is what always worries me about our home server. It's running ZFS with multiple redundant drives but the supplier refused (when I explicitly asked) to supply it with disks known to be from different batches claiming that the odds of multiple failures close together were negligible. Obviously we have backups as well but the time and cost to restore a full server from online backups can be significant.

My solution is to use a different manufacturer for each drive in a mirror. The prices are usually pretty similar and you get to make sure that one firmware bug doesn't kill your entire pool.

I decided to employ this tactic when I was setting up my new NAS and needed two drives.

Upside was that I could definitely know they weren't from the same batch.

Downside was that I had to buy a Seagate, and I don't have good experiences with Seagate since my only Seagate drive had died an early death at the tender age of 3. Turns out that this was very much a downside since the Seagate drive died at the tender age of 16 months.

Re: SSD will fail at 40k power-on hours (2021)

#162

Yikes. Cisco claims this is "an industry wide firmware index bug". Is there any validity to this claim? Are any consumer drives affected by the same issue?

Yes, HPE is one of the SSD OEMs affected by it:

https://support.hpe.com/hpesc/public/docDisplay?docId=emr_na...

https://support.hpe.com/hpesc/public/docDisplay?docLocale=en...

> Are any consumer drives affected by the same issue?

As far as I know it doesn't affect consumer drives, but I wouldn't be surprised if some have the same defective firmware.

Re: SSD will fail at 40k power-on hours (2021)

#163
post #161
post #139

Earlier quoted context omitted.

My solution is to use a different manufacturer for each drive in a mirror. The prices are usually pretty similar and you get to make sure that one firmware bug doesn't kill your entire pool.

I decided to employ this tactic when I was setting up my new NAS and needed two drives. Upside was that I could definitely know they weren't from the same batch. Downside was that I had to buy a Seagate, and I don't have good experiences with Seagate since my only Seagate drive had died an early death at the tender age of 3. Turns out that this was very much a downside since the Seagate drive died at the tender age o…

I had a series of WD drives that failed, and I managed to get them all replaced under warranty since they died within 3-5 years. I don't buy WD drives anymore but it wasn't the end of the world since I had spares while I waited for the replacement drives to be shipped.

Anecdotally, I haven't had issues with Seagate but I'm sure it really boils down to which exact drives you're using and what batch they were in.

Re: SSD will fail at 40k power-on hours (2021)

#164
post #139

Earlier quoted context omitted.

This is what always worries me about our home server. It's running ZFS with multiple redundant drives but the supplier refused (when I explicitly asked) to supply it with disks known to be from different batches claiming that the odds of multiple failures close together were negligible. Obviously we have backups as well but the time and cost to restore a full server from online backups can be significant.

My solution is to use a different manufacturer for each drive in a mirror. The prices are usually pretty similar and you get to make sure that one firmware bug doesn't kill your entire pool.

I do the same. Ages ago I once had to build a server at work and picked three vendors for the raid 5. Got funny looks drom coworkers, apparently they found the idea super strange. One drive (Seagate of course) failed after a year, and since we had a matching size WD lying around, used that. Now there were two WDs in the setup. Some years later the PSU blew up and killed both WDs, the Toshiba survived.

Re: SSD will fail at 40k power-on hours (2021)

#165

Earlier quoted context omitted.

Quoted post unavailable.

Barry I appreciate that you recognize my good faith efforts. But I want to highlight that queer and gender nonconforming people are regularly marginalized and othered in this country and around the world. It is genuinely tiring to them to be dismissed regularly in their daily life and then to encounter people online who want to play this up as some culture war with two legitimate sides. I personally would not label y…

No post body was provided.

Re: SSD will fail at 40k power-on hours (2021)

#166
post #134

Earlier quoted context omitted.

Yowch. The old "stagger your drive replacements, stagger your batches" thing might not be quite as outdated as we'd like to think...

I had no idea anyone thought this was outdated. We certainly never stopped doing it. I think it is a timeless failsafe.

Yeah, Backblaze and DigitalOcean both talk about it a bunch in their sysops stuff.

Re: SSD will fail at 40k power-on hours (2021)

#167
post #152

Earlier quoted context omitted.

Within an earlier thread "Tell HN: HN Moved from M5 to AWS", there's an excellent comment by loxias about risk diversification across multiple factors. Well worth reading: https://news.ycombinator.com/item?id=32031655 I've increasingly come to view systems operations / SRE as a risk management exercise, where the goal is to reduce the odds of a catastrophic failure. Total system outage is one level, unrecoverable tot…

> I've increasingly come to view systems operations / SRE as a risk management exercise, where the goal is to reduce the odds of a catastrophic failure. s/reduce/find an appropriate level for/ It's a common misconception that risk management and risk reduction are synonyms. Risk management is about finding the right level of risk given external factors. Sometimes that means maintaining the current level of risk or ev…

My point is that risk is central to systems management. If you look at earlier standard texts on the subject, e.g., Nemeth or Frisch, the concept of risk is all but entirely missing. I've numberous disagreements with Google, but one place where I agree is that the term SRE, systems reliability engineer, puts the notion of managing for stability front and centre, and inherently acknowledges the principle of risk. I've since heard from others that this is in fact how the practice is presented and taught there.

Quibbling over whether the proper term is risk management or risk reduction rather spectacularly misses the forest for the trees.

Re: SSD will fail at 40k power-on hours (2021)

#168
post #78
post #30

Earlier quoted context omitted.

This will not affect your laptop, all of the models affected by this are enterprise SAS SSDs. Of course your SSD might have some other firmware bug that would eat your data, all you can do is search for the model number and see if the manufacturer has issued any notices/firmware updates.

There was at least one consumer SSD with a similar failure mode, the Crucial M4 SATA drive, unless you updated the firmware it would crash after 5200 cumulative power on hours. That drive launched in 2011 though so there probably aren't that many still in active use which still haven't reached ~7 months of uptime.

Yes. This one: https://www.reddit.com/r/buildapc/comments/1z2rm5/crucial_m4...

That problem became known a decade ago, so it's somewhat surprising to see such a similar bug now.

This new one is worse because the drive cannot be used after reaching the magic number of hours. In the Crucial M4 case the firmware could be updated even after the bug struck.

Re: SSD will fail at 40k power-on hours (2021)

#169

Yikes. Cisco claims this is "an industry wide firmware index bug". Is there any validity to this claim? Are any consumer drives affected by the same issue?

Is there any information about the provenance of this SSD controller? Sounds like enterprisey venodrs all rebranded some upstream supplier's hardware.

edit: apparently sandisk: http://forum.hddguru.com/viewtopic.php?f=3&t=39964 - also a clue about the magic 40k hour significance: "the SSD alters its performance in some way as it approaches end of life. This appears to shine some light on the reason for a trigger at 40K Power On Hours."

Re: SSD will fail at 40k power-on hours (2021)

#170

Earlier quoted context omitted.

Why would there be submissions? These drives predate HN.

HN occasionally discusses issues pre-dating itself.

... and the lifetime (deathtime?) award for the Deathstar only predated HN by a few months:

May 26, 2006: https://www.pcworld.com/article/535838/worst_products_ever.h...

October 9, 2006: https://news.ycombinator.com/item?id=1

Memory would still have been reasonably green.

Post reply on HN