Live data from Hacker News

SSD will fail at 40k power-on hours (2021)

cisco.com

81–90 of 276 posts

Re: SSD will fail at 40k power-on hours (2021)

#81
post #74

Earlier quoted context omitted.

Checked my Samsung 970 Evo 2TB and it says 487 even though it’s been on continuously for years.

Receiving power isn’t the same as being “on”. I’d assume the drive has a sleep state that it goes into after inactivity, and those hours don’t count as “power on” hours.

Interesting, thanks. The spec does leave it almost uselessly underspecified:

"Power On Hours: Contains the number of power-on hours. This may not include time that the controller was powered and in a non-operational power state."

The same drive reports only 329 "controller busy time" minutes.

Re: SSD will fail at 40k power-on hours (2021)

#82
post #28

Check your power-on hours: $ sudo smartctl -a /dev/sda | grep -e Power_On_Hours -e ^ID ID# ATTRIBUTE_NAME FLAG VALUE WORST THRESH TYPE UPDATED WHEN_FAILED RAW_VALUE 9 Power_On_Hours 0x0032 098 098 000 Old_age Always - 9743 Just looking at the raw value, it seems to be 9'743 hours in my case

iMac mid-2010. Original disk drive.

"9 Power_On_Hours 0x0032 001 001 000 Old_age Always - 74233"

12 years old. More than 8 years of run time. It keeps on purring.

Yes, I have redundant backups. I also have a replacement drive ready. I just want to see how far I can take it.

Re: SSD will fail at 40k power-on hours (2021)

#83

Earlier quoted context omitted.

Quoted post unavailable.

If someone thinks inclusive language is a sign of hostility, then I think they have misunderstood the situation. I’m happy to have open source contributions from anyone, but I do share my pronouns in posts and on videos and I can tell you that the very small number of people who have gotten upset by that were angry unhelpful people who were more interested in complaining about language than contributing to the commun…

No post body was provided.

Re: SSD will fail at 40k power-on hours (2021)

#84
post #48
post #30

Earlier quoted context omitted.

This will not affect your laptop, all of the models affected by this are enterprise SAS SSDs. Of course your SSD might have some other firmware bug that would eat your data, all you can do is search for the model number and see if the manufacturer has issued any notices/firmware updates.

> This will not affect your laptop That’s just your presumptive opinion, right? Edit: sorry, probably put that offensively. mikiem said about the HN drives: “These were made by SanDisk (SanDisk Optimus Lightning II) and the number of hours is between 39,984 and 40,032...” - https://news.ycombinator.com/item?id=32031428 Without knowing parts of a codebase are shared between SanDisk devices, it is hard to say that ente…

I've been searching "40000 hour SSD" since the HN downtime. There's a lot of bug reports besides this one and I'm fairly confident it only affects enterprise too.

Re: SSD will fail at 40k power-on hours (2021)

#85
post #68

Somewhat unrelated, but I recently had a motherboard fried by power instability which has given me a healthy respect for the difference between spinning rust and SSD's. My SSD's were scragged. my HDD's were just fine. I guess it's time to figure out how to get a realistic write-through cache setup going, because from now on, if it ain't on spinnin' rust it ain't hard enough yet.

Of what quality was the mobo? I thought the higher-end stuff typically has power protection.

Re: SSD will fail at 40k power-on hours (2021)

#86
post #9
post #7

Earlier quoted context omitted.

Wow, thanks for sharing. I didn't realize how closely related they were. (TLDR For anyone wondering, "recent HN issues" means HN very likely went down yesterday because of this same bug, when two (edit: two pairs, four total) enterprise SSDs with old firmware died after 40,000 hours close together. An admin of HN and its host both like this theory. See details in that thread.) Edit: If you want to discuss that theory…

the chance of two SSD's failing at the same time under normal circumstances is extremely slim. So this might actually be a good cause of this incident.

If both SSD's are from the same lot number and one fails, the chances of the second failing go up by a high amount. Both failing at the same time though is extremely rare.

Re: SSD will fail at 40k power-on hours (2021)

#87
post #9
post #7

Earlier quoted context omitted.

Wow, thanks for sharing. I didn't realize how closely related they were. (TLDR For anyone wondering, "recent HN issues" means HN very likely went down yesterday because of this same bug, when two (edit: two pairs, four total) enterprise SSDs with old firmware died after 40,000 hours close together. An admin of HN and its host both like this theory. See details in that thread.) Edit: If you want to discuss that theory…

the chance of two SSD's failing at the same time under normal circumstances is extremely slim. So this might actually be a good cause of this incident.

It seems more likely it was four drives (though dang and Mike both refer to "two" in the earlier thread).

Both primary and failover servers had RAID arrays. I suspect RAID 10 (striped mirror), which would mean two drives would have to fail to take down a single server.

Four drives of the same manufacturer spec and batch would do that.

Re: SSD will fail at 40k power-on hours (2021)

#88
post #28

Check your power-on hours: $ sudo smartctl -a /dev/sda | grep -e Power_On_Hours -e ^ID ID# ATTRIBUTE_NAME FLAG VALUE WORST THRESH TYPE UPDATED WHEN_FAILED RAW_VALUE 9 Power_On_Hours 0x0032 098 098 000 Old_age Always - 9743 Just looking at the raw value, it seems to be 9'743 hours in my case

    ID# ATTRIBUTE_NAME          FLAG     VALUE WORST THRESH TYPE      UPDATED  WHEN_FAILED RAW_VALUE
    9 Power_On_Hours          0x0032   055   055   000    Old_age   Always       -       39676
I'm 300 hours from 40K, time to buy new SSD? is this real?!

Re: SSD will fail at 40k power-on hours (2021)

#89
Would companies be willing to contribute to the OpenSSD project?

OCP (Open Compute Project) has shown that customer-operators can cooperate on open hardware designs, successfully influencing enterprise hardware supply chains. Commercial DPUs and SmartNICs were preceded by a decade of open hardware and research by the NetFPGA project (https://netfpga.org). Why not DiskFPGA?

2017 OpenSSD overview, based on Xilinx: https://github.com/Cosmos-OpenSSD/Cosmos-plus-OpenSSD/blob/m...

2022 status, http://www.openssd-project.org/

> OpenSSD platforms are still being actively used in many academic institutions. As of June 2022, we have renewed the homepage hoping that this site will be a forum to share various simulators, tools, traces, etc. not only for the conventional SSDs but also for the upcoming storage devices such as KVSSD, ZNS SSD, and Computational Storage (CSX). This site is being maintained by Systems Software and Architecture Lab. at Seoul National University as a part of the SW STAR Lab. project.

Re: SSD will fail at 40k power-on hours (2021)

#90
post #28

Check your power-on hours: $ sudo smartctl -a /dev/sda | grep -e Power_On_Hours -e ^ID ID# ATTRIBUTE_NAME FLAG VALUE WORST THRESH TYPE UPDATED WHEN_FAILED RAW_VALUE 9 Power_On_Hours 0x0032 098 098 000 Old_age Always - 9743 Just looking at the raw value, it seems to be 9'743 hours in my case

ID# ATTRIBUTE_NAME FLAG VALUE WORST THRESH TYPE UPDATED WHEN_FAILED RAW_VALUE 9 Power_On_Hours 0x0032 055 055 000 Old_age Always - 39676 I'm 300 hours from 40K, time to buy new SSD? is this real?!

I mean, is it the affected model?

Have you applied the appropriate firmware update?

Post reply on HN