Live data from Hacker News

Why I'm usually unnerved when modern SSDs die on us

utcc.utoronto.ca

211–220 of 258 posts

Re: Why I'm usually unnerved when modern SSDs die on us

#211

Earlier quoted context omitted.

I've noticed the same thing, and suspect that it's related to the way that higher level software scales compared to embedded. If a line of code is written to run in a customer's browser, then that line of code may be deployed to millions, maybe billions, of customers. But, if an equivalent line of code goes in to firmware for some widget, then you're doing pretty good to get that line in to a million widgets at all,…

Write an embedded bootloader and there's a good chance your code will be used by a billion people within a few years of writing it. You'll still get paid peanuts for doing it.

I think it's a issue related to the visibility of the quality of the work. If the product is even 20% more reliable (whatever that exactly means), hardly anyone will actually notice. Nobody notices the absence of an error. That's although it probably took a huge effort to achieve this. Making a user flow just a little bit nicer is very visible and gets attention.

I always felt that ISPs suffer a similar problem. Nobody cares if everything works as expected, but we'll get upset of it doesn't. However, there is hardly anything outside of what we take for granted they can do that we will actively appreciate. What could a embedded engineer working on SSDs do that will be noticed, appreciated and not taking for granted by customers?

Re: Why I'm usually unnerved when modern SSDs die on us

#212

Earlier quoted context omitted.

i've never seen or heard of an automated board rework system. thinking about the steps and things i've had to do to manually rework boards, your machine would need not only be able to apply force to pull parts off of boards without damaging the board, but also be prepared to restore pads / through holes to usable states after desoldering parts, before new ones could go back in. there's a reason companies don't repair…

I've done it myself, save restoring pads -- to me it's just the sort of fiddly, finnickety thing that it seems like robots sold be good at. And, they're not going to burn their fingers or run out of hands to hold things!

Robots are good at repetitive tasks and economy of scale. Doing a tricky action for ten thousand boards is something that robots are good at; doing a different tricky action on each of a hundred boards is not.

It doesn't seem plausible to have a business case where you'd get be able to get a large quantity of identical boards (that are otherwise good!), replace caps on them, and be able to sell them for much more than you got them - i.e. that the boards haven't become obsolete in that time. If there's no mass production, there's not much use for automation.

Re: Why I'm usually unnerved when modern SSDs die on us

#213
My first experience with drive failure was a ~40MB HDD expansion card in a 386. The bearings got "sticky", so the spindle wouldn't start rotating. But there was a Al tape covered hole, and you could insert the eraser end of a pencil, and nudge it. So yes, very understandable.

Not too much later, I used Iomega ZIP drives, and experienced the "click of death". That was sudden, and irreversible, but also very understandable.

For the past couple decades, I've consistently used RAID arrays, mostly RAID1 or RAID10 (and RAID0 or RAID5-6 for ephemeral stuff). I've had several HDD failures, but they were usually progressive, and I just swapped out and rebuilt.

I recently had my first SSD failure. And it was also progressive. The first symptom was system freeze, requiring hard reboot, and then I'd see that one of the SSDs had dropped out of the array. But I could add it back. At first, I thought that there was some software problem, and that the RAID stuff was just caused by hard reboot.

But eventually, the box wouldn't boot, so I had to replace the bad SSD and rebuilt the array. It was complicated by having sd1 RAID10 for /boot, and sd5 RAID10 for LVM2 and LUKS. So I also had to run fdisk before device mapper would work.

Re: Why I'm usually unnerved when modern SSDs die on us

#214
post #25

Maybe I am being overly simplistic, but shouldn't it not matter? Who in the modern age doesn't back up everything all the time? Don't we all operate with the assumption these things are going to blow at any time? 90%+ of my data is on cloud storage now anyway. When a SSD goes out don't you just chunk it in the drawer of old drives that you promise to take to the disposal center this weekend (and never do) and then ta…

Replacing an SSD is not free, and in most cases it's not easy. Maybe an IT pro can just roll down to the computer store for a new one, put it in their laptop (for free!), and throw a $100+ drive in a drawer without even thinking about warranty, but most people can't. A backup doesn't excuse excessive rates of failure and weird glitches.

The post is about predictability and warnings before a drive dies. The work and cost to replace it doesn't change anyway (since you do replace the drive at the first sign of warnings, right? otherwise what's the point of wanting them?). The only difference is if you get an extra chance to copy the data before you replace the drive - which is no difference at all if you have proper backups.

If we'd be arguing about mean time between failure or total cost of ownership, then it'd be relevant, but this post isn't even claiming that the rates of failure are excessive (compared to what?), just that they are too weird and unpredictable for the author's liking.

Re: Why I'm usually unnerved when modern SSDs die on us

#215

Earlier quoted context omitted.

You still get paid a lot more working at google working on generic backend protobuf shuffling than you will working on SSD firmware at a hardware company or intel's C++ compiler.

For those doubting you, the going rate for embedded engineers out here in the Denver area where a lot of these SSD controllers are designed is ~$90k. Embedded engineers get peanuts for some reason.

In Xbox, many of the firmware hires the hardware folks made seemed to be payed poorly. Also, there were tons of contractors, and not much institutional knowledge was retained. (At one point they had to pay a consulting firm to decompile the firmware for a controller because they had lost the source code).

"Why do we need source control? It's all there, right on my laptop. Source Depot is just a bunch of trouble." [rough quote from memory, maybe conflated from a couple of engineers]. I'm happy to report that things got better.

On the software side of Xbox the people were much better compensated, and we wrote lots of firmware, too. It was probably harder to be hired, though.

Re: Why I'm usually unnerved when modern SSDs die on us

#216
post #215

Earlier quoted context omitted.

For those doubting you, the going rate for embedded engineers out here in the Denver area where a lot of these SSD controllers are designed is ~$90k. Embedded engineers get peanuts for some reason.

In Xbox, many of the firmware hires the hardware folks made seemed to be payed poorly. Also, there were tons of contractors, and not much institutional knowledge was retained. (At one point they had to pay a consulting firm to decompile the firmware for a controller because they had lost the source code). "Why do we need source control? It's all there, right on my laptop. Source Depot is just a bunch of trouble." [ro…

"Why do we need source control? It's all there, right on my laptop. Source Depot is just a bunch of trouble."

I've threatened to withhold paychecks from employees who have said this to me. The job isn't done till the code is checked in, building, and backed up.

Re: Why I'm usually unnerved when modern SSDs die on us

#217

Earlier quoted context omitted.

For those doubting you, the going rate for embedded engineers out here in the Denver area where a lot of these SSD controllers are designed is ~$90k. Embedded engineers get peanuts for some reason.

Weird, I thought embedded would be much more sought after and rare.

There's A LOT less people. But the demand is also lower than for people than can build simple websites. It seems like the latter is a lot more important than the first for setting the market rate on salaries. Or maybe it's about the fact that embedded work is often done be electrical engineers, which are seen in a different pay category in some countries.

Re: Why I'm usually unnerved when modern SSDs die on us

#218

Earlier quoted context omitted.

You still get paid a lot more working at google working on generic backend protobuf shuffling than you will working on SSD firmware at a hardware company or intel's C++ compiler.

For those doubting you, the going rate for embedded engineers out here in the Denver area where a lot of these SSD controllers are designed is ~$90k. Embedded engineers get peanuts for some reason.

I know what you mean--they could get paid a lot more elsewhere--but it still weirds me out that techies consider $90k/year "peanuts". I know people trying to raise children on one-third of that.

Re: Why I'm usually unnerved when modern SSDs die on us

#219
post #91

I worked on SSD firmware for quite a long time and here is my perspective. Early flash used to fairly reliable with almost minimal error correction. However with increasing density, smaller processes and multi level cells, it has gone progressively less reliable and slower. Here are some of the things that we need to worry about: https://www.flashmemorysummit.com/English/Collaterals/Procee... To compensate for all th…

You can even offer 0 equity, but pretty sure if you offer a $180k salary you'll have a flock of new grads apply.

The reason you're not getting enough new grads is probably because your compensation isn't competitive enough.

Re: Why I'm usually unnerved when modern SSDs die on us

#220

I've experienced a few seriously strange issues with modern SSDs, even some of the better ones. I had a 512GB Samsung drive that became very slow randomly at doing IO operations, the whole machine would die for 10-30 seconds at a time once or twice a day while any process that tried to use the disk became blocked on IO. Then it'd come right back like everything was perfectly fine. Issues like this definitely worry me…

If on Linux, try running fstrim from time to time. More free space makes life of SSD's garbage collector / defragmenter much easier. I've anecdotally noticed that running fstrim reduces freezeups under heavy load from 1-2s to almost nothing on my Toshiba drive.

On Windows 10, I believe it's supposed to automatically TRIM free space on NTFS.
Post reply on HN