Live data from Hacker News

Why I'm usually unnerved when modern SSDs die on us

utcc.utoronto.ca

161–170 of 258 posts

Re: Why I'm usually unnerved when modern SSDs die on us

#161

Earlier quoted context omitted.

The claim was hundreds of picometers, which is an order of magnitude smaller than a few nanometers. Literally nothing can be perfectly flat, nor perfectly aligned.

It doesn't need to be. The basic idea is that the fly height is self-regulating; if the head goes away from the platter, its "lift" is reduced, so the springiness of the arm forces it back to the platter. If it moves closer to the platter, lift is increased, so it moves away. Similarly the tracks don't have to be perfectly round or concentric, because the lowest-level head control system in the drive actively tracks…

> Similarly the tracks don't have to be perfectly round or concentric

They have to be pretty dang close at 7200+ rpm.

Re: Why I'm usually unnerved when modern SSDs die on us

#162
post #104

Earlier quoted context omitted.

I'm mostly curious because I work in a storage-adjacent field (NAS) for a BigCorp and the pay is pretty good, if not quite FAANG level. It's not a startup by any means, but I will easily become a multi-millionaire in a handful of years. I was curious about the other side of the fence.

Huh? If you're going to "easily become a multi-millionaire in a handful of years", then your pay is more than pretty good, and certainly not worse than FAANG level.

Sorry, handful of years from today. I've been working for 7 years now. FAANG comp would probably be 20%-25% higher; I'm mostly good at keeping my expenses down and saving a high proportion of my income. I've also had the good fortune of the bull market working in my favor for the entire time I've been employed.

Re: Why I'm usually unnerved when modern SSDs die on us

#163
post #104

Earlier quoted context omitted.

I'm mostly curious because I work in a storage-adjacent field (NAS) for a BigCorp and the pay is pretty good, if not quite FAANG level. It's not a startup by any means, but I will easily become a multi-millionaire in a handful of years. I was curious about the other side of the fence.

What is a handful of years? That sounds like FAANG pay to me.

Sorry, handful of years from today. I've been working for 7 years now. FAANG comp would probably be 20%-25% higher; I'm mostly good at keeping my expenses down and saving a high proportion of my income. I've also had the good fortune of the bull market working in my favor for the entire time I've been employed.

Re: Why I'm usually unnerved when modern SSDs die on us

#164
post #91

I worked on SSD firmware for quite a long time and here is my perspective. Early flash used to fairly reliable with almost minimal error correction. However with increasing density, smaller processes and multi level cells, it has gone progressively less reliable and slower. Here are some of the things that we need to worry about: https://www.flashmemorysummit.com/English/Collaterals/Procee... To compensate for all th…

Let's be honest, the chance of a multi-million dollar payoff for an employee is low regardless. Young entry-level devs are optimistic and also vulnerable to believing a line of bullshit on how much their options might be worth some day. I do agree it is a "higher" chance in web/mobile technology, sort of like how your chance of winning the lottery is "higher" if you buy 10 tickets instead of 1.

You still get paid a lot more working at google working on generic backend protobuf shuffling than you will working on SSD firmware at a hardware company or intel's C++ compiler.

Re: Why I'm usually unnerved when modern SSDs die on us

#165

Earlier quoted context omitted.

Huh? If you're going to "easily become a multi-millionaire in a handful of years", then your pay is more than pretty good, and certainly not worse than FAANG level.

Maybe his hands are larger than ours ;)

Only the ten digits ;-)

Re: Why I'm usually unnerved when modern SSDs die on us

#166
post #93

Earlier quoted context omitted.

Users and administrators almost certainly prefer a 20 minute IO latency over data corruption. Host operating systems should probably flag an IO as failed long before 20 minutes and then you know 1) nothing made it to disk and 2) have some chance to avoid introducing additional corruption, if e.g., the OS is smart enough to kick out the drive when this happens. > Another problem is that the job while rewarding is not…

> Users and administrators almost certainly prefer a 20 minute IO latency over data corruption. If the drive part of a RAID setup I would actually prefer it just reports itself failed and doesn't slow down access to the array by scanning itself for 20 minutes. As far as I know, that is one of the main differences when buying enterprise or nas drives compared to consumer drives. With nas drives, the firmware gives up…

Potato, potato. If your array doesn't kick out a drive that stalls for 30 seconds, much less 20 minutes, why not?

Re: Why I'm usually unnerved when modern SSDs die on us

#167
post #50

Earlier quoted context omitted.

I have seen some weird issues with SSDs. I had an OCZ Vertex 2 die on me multiple times, but one thing that stood out most is that after a power cycle or complete system shutdown (note: reboots were just fine) everything I had done last time - install software, update Windows, create files - was gone. The state was reverted to before the time it booted. It was like my computer contained some kind of Reborn chip, exce…

Like many early adopters, I too had a bunch of failed Vertex 2 drives and sometimes observed similar things. I think this might be because the drive lost some updates to its FTL about where it wrote new data, which could plausibly lead to both new files vanishing and changes to existing ones being apparently undone.

Yes, the Sandforce controller in Vertex 2 was a real unstable beast. I lost maybe three or four Vertex 2 drives. All under warranty, except the last one which suddenly vanished and was not detected anymore.

OCZ does not exist anymore. Not sure of if this was one of the causes, but I would otherwise never buy a drive from them again.

Re: Why I'm usually unnerved when modern SSDs die on us

#168
post #91

I worked on SSD firmware for quite a long time and here is my perspective. Early flash used to fairly reliable with almost minimal error correction. However with increasing density, smaller processes and multi level cells, it has gone progressively less reliable and slower. Here are some of the things that we need to worry about: https://www.flashmemorysummit.com/English/Collaterals/Procee... To compensate for all th…

Are you saying that TLC SSDs are essentially unreliable and it is a miracle we don't see higher failure rates?

Regarding your point regarding languages: I would be interested and motivated, but I probably don't have the skill (probably, only did some x86 ASM/C++) nor the location (Europe). Usually someone doesn't just start with C++ but with a managed language, and once someone lands a job becomes demotivated or just doesn't have enough time.

Re: Why I'm usually unnerved when modern SSDs die on us

#169
post #41

"When a HD died early, you could also imagine undetected manufacturing flaws that finally gave way. With SSDs, at least in theory that shouldn't happen" Why shouldn't it? Isn't it just hardware too? "With spinning HDs, drives might die abruptly but you could at least construct narratives about what could have happened to do that" Why can't you do the same with SSDs? It feels like the author's main complaint is the fr…

"It feels like the author's main complaint is the frustration of not understanding SSD hardware as well." What is so frustrating about SSDs is how very poorly they compare to previous incarnations of solid state storage. Using Disk-On-Chip and/or IDE-pin-compatible CF cards, I had many, many devices in the field that lasted, mounted read-only, for decades An entire sect of the computing industry came to rely on these…

Every failure of an SSD feels like an exceptional event. Some harbinger of doom that needs to be shouted from the rooftops. The prions of storage.

But magnetic hard drives failed all the time. I have a giant stack in my office closet just from my dev machines over the years. But it wasn't new and scary -- it was just a hard drive failing -- so it was just normal. Some had controllers fail, suddenly blinking out of existence. Another had cache memory corrupt so it just gave ridiculous readings occasionally. Others had physical failures.

I don't know where to begin relative to prior flash memory (e.g. CF cards) which were absolutely notorious trash.

It is worth noting that every smartphone the world over has an "SSD" in it. We spend remarkably little of our mental power concerned about the flash storage. It is the cause of a negligible amount of device failures.

Re: Why I'm usually unnerved when modern SSDs die on us

#170
post #122

Earlier quoted context omitted.

> Even your compiled machine language is ultimately far more abstracted from what the processor actually does than it was on, say, a 6502 Funny that you bring up 6502; that reminds me of 1541 disk drive for C64, which had mostly same 6502 as the host computer (albeit running at slower speed).

Most disk drives of the time were like that. One of the reasons the Apple II disk drive was so affordable is that Woz just used the Apple II's own 6502 to handle the grunt work, with their disk drive being little more than a drive mechanism, a PROM, and some ICs. Since this is The Woz we're talking about, he went ahead and broke with conventional encoding while he was at it and instead implemented a scheme which allo…

The Apple ][ disk encoding was not done to allow "for a few extra sectors per track". Since the Apple ][ did not have a disk controller, the encoding from magnetic flux to bits was handled in software. A 1Mhz 6502 cannot handle the "standard" encoding (things like M2FM) in software as the disc was rotating. Woz's encoding allowed the encoding.
Post reply on HN