Live data from Hacker News

Why I'm usually unnerved when modern SSDs die on us

utcc.utoronto.ca

181–190 of 258 posts

Re: Why I'm usually unnerved when modern SSDs die on us

#181

Earlier quoted context omitted.

I don't know if I'm getting exactly what you're saying, but a sudden power off isn't going to wipe your SSD's SLC cache like it would do the DRAM cache on more expensive drives.

But if the controller gets into a weird state and "forgets" to flush the cache and then doesn't even look at the SLC cache on the next boot...

The SLC cache is managed by the controller, it's not going to forget to write it to the flash cells being managed as TLC/QLC, the controller is transparently storing your data in both. The controller knows if put some data in the SLC cache and some data in the QLC, and that that data needs to move from SLC to QLC, and that that other data in the SLC was marked deleted so now it can return that entire block to being addressed as QLC.

I don't know. I'm just trying to say that since flash isn't volatile the SLC cache isn't treated as volatile cache either.

Re: Why I'm usually unnerved when modern SSDs die on us

#182

Earlier quoted context omitted.

> Another problem is that the job while rewarding is not very lucrative. The chance of a multi million dollar payoff for an employee is low. I have a higher chance working on a web connected gadget to become a millionaire. So that means it is really hard to recruit those who are top notch programmers who known how to figure out the algorithms, write the code, debug the hardware. Most new grads these days are interest…

I don't know how the system could / should interface with stupid trends sucking resources (high complexity projects, smart individuals, long term).

Infrastructure almost always becomes a commodity, or at least moves behind the scenes. For it to be lucrative, interesting technical position, you have to have a large-scale enterprise (like Google), serving millions of people, and you have to promote it as a worthwhile career path (also done at Google as the SRE position).

Re: Why I'm usually unnerved when modern SSDs die on us

#183
I've done a bit of ad-hoc reliability testing with SSD's.

Some years ago I got a great deal on several Pacer disks and wrote a program to write a pseudo-random sequence of data (using a known initial seed) across the entire disk and read it back and compare. Part way through, the data didn't match. No ECC errors, nothing raised by the filesystem, just mismatched bits which came back in a manner which tried to "trick" me into thinking they were good data. This happened on like 5 of the 8 disks. Needless to say I sent those crappy SSD's back to the manufacturer (unfortunately only got a 2/3 refund) along with some harsh words for their engineers.

I've had more name-brand SSD's fail, in various manners (even on well-reviewed Kingston drives). Sometimes in ways that can't be accessed at all, other times (at best of times) in a manner which doesn't allow writes but still allows reads (albeit at a trickle of a datarate).

These days I use solely Intel-based, top-line SSD's, and some (very limited) Samsungs. The choice isn't based on empirical data, but rather and impression their bar is a little higher (or more conservative) in terms of reliability, and simply not wanting to deal with the apparent issues I've seemed to encounter with other brands. The downtime lost from restoring / reconstructing just isn't worth it to me. Maybe I'm paying twice as much as I ought to, but since making the switch many years back it's worked out pretty well and I've been happy / fortunate.

I run my SSD's in RAID10 using high-end controllers (aside from a few in ZFS).

Just my own subjective experiences, again I'm not doing this at scale.

Re: Why I'm usually unnerved when modern SSDs die on us

#184

Earlier quoted context omitted.

Let's be honest, the chance of a multi-million dollar payoff for an employee is low regardless. Young entry-level devs are optimistic and also vulnerable to believing a line of bullshit on how much their options might be worth some day. I do agree it is a "higher" chance in web/mobile technology, sort of like how your chance of winning the lottery is "higher" if you buy 10 tickets instead of 1.

You still get paid a lot more working at google working on generic backend protobuf shuffling than you will working on SSD firmware at a hardware company or intel's C++ compiler.

For those doubting you, the going rate for embedded engineers out here in the Denver area where a lot of these SSD controllers are designed is ~$90k. Embedded engineers get peanuts for some reason.

Re: Why I'm usually unnerved when modern SSDs die on us

#185
post #12

"We had one SSD fail in this way and then come back when it was pulled out and reinserted, apparently perfectly healthy, which doesn't inspire confidence." We've experienced exactly the same thing. Our general course of action is to perform a hard power cycle of the server through IPMI - a warm cycle doesn't seem to work. I've always presumed it was down to dodgy SSD controller firmware given the way it suddenly stop…

I have three SSDs in three different laptops/desktops. In their current host machines they've been working flawlessly for a couple of years. Prior to my figuring out which SSD paired best with which host machine, I experienced intermittent strange and catastrophic problems (unreadable sectors to complete data loss) with each one. These were different brands, different capacities, bought in different years. It's sort…

I just don't keep anything important on my SSD. My desktop's SSD is for Windows and games. All documents and other stuff goes on my mechanical drives and my AppData folder is backed up every night too. Everything on my laptop's SSD is either in cloud storage or in an external git repo. I'm 100% prepared for the certain eventuality of any of these SSDs going tits up unexpectedly and catastrophically.

I'm more worried about everybody else who gets an SSD and doesn't take the right precautions, because everybody sells SSDs as being so much more reliable than mechanical drives.

I worked as a PC technician for a while recently. Of the handful of catastrophic failures of mechanical drives we had, the majority of those were ones that were physically dropped, resulting in a head crash. Otherwise, we generally always managed to save data from failing drives. Any failing SSD we encountered was just dead, since there are really only two states: Fine or failed. There was nothing we could do except refer them to a data recovery company that charges thousands of Euros.

Re: Why I'm usually unnerved when modern SSDs die on us

#186

Earlier quoted context omitted.

You still get paid a lot more working at google working on generic backend protobuf shuffling than you will working on SSD firmware at a hardware company or intel's C++ compiler.

For those doubting you, the going rate for embedded engineers out here in the Denver area where a lot of these SSD controllers are designed is ~$90k. Embedded engineers get peanuts for some reason.

~$90k is now peanuts. Interesting.

Re: Why I'm usually unnerved when modern SSDs die on us

#187

Earlier quoted context omitted.

You still get paid a lot more working at google working on generic backend protobuf shuffling than you will working on SSD firmware at a hardware company or intel's C++ compiler.

For those doubting you, the going rate for embedded engineers out here in the Denver area where a lot of these SSD controllers are designed is ~$90k. Embedded engineers get peanuts for some reason.

Weird, I thought embedded would be much more sought after and rare.

Re: Why I'm usually unnerved when modern SSDs die on us

#188
post #93

Earlier quoted context omitted.

Users and administrators almost certainly prefer a 20 minute IO latency over data corruption. Host operating systems should probably flag an IO as failed long before 20 minutes and then you know 1) nothing made it to disk and 2) have some chance to avoid introducing additional corruption, if e.g., the OS is smart enough to kick out the drive when this happens. > Another problem is that the job while rewarding is not…

As someone that identified a bug in a Drobo firmware once and was offered a job on the spot I think the problem with attracting talent is two fold. The first problem is really two parts, not only is it rare to find people who have passion for storage related technologies but very few will gain exposure to these technologies to develop that passion. Kids don't routinely grow up with a SAN in the house. They do tend to…

>As someone that identified a bug in a Drobo firmware once and was offered a job on the spot

Just curious how you did that. Did you disassemble their firmware?

Re: Why I'm usually unnerved when modern SSDs die on us

#189

Earlier quoted context omitted.

For those doubting you, the going rate for embedded engineers out here in the Denver area where a lot of these SSD controllers are designed is ~$90k. Embedded engineers get peanuts for some reason.

~$90k is now peanuts. Interesting.

That's less than 30k in 1980 dollars. Yeah, that's next to nothing.

Re: Why I'm usually unnerved when modern SSDs die on us

#190

Earlier quoted context omitted.

You still get paid a lot more working at google working on generic backend protobuf shuffling than you will working on SSD firmware at a hardware company or intel's C++ compiler.

For those doubting you, the going rate for embedded engineers out here in the Denver area where a lot of these SSD controllers are designed is ~$90k. Embedded engineers get peanuts for some reason.

Can confirm former Intel SSD in longmont 95k -- shit RSUs.
Post reply on HN