Live data from Hacker News

HN is up again

news.ycombinator.com

361–370 of 390 posts

Re: HN is up again

#361
post #307

Earlier quoted context omitted.

You are never going to guess how long the HN SSDs were in the servers... never ever... OK... I'll tell you: 4.5years. I am not even kidding.

How many other customers will/have hit this?

Every large DC will have hit it (Amazon, Facebook, Google, etc). But it's a shame that all their operational knowledge is kept secret.

Re: HN is up again

#362

Earlier quoted context omitted.

> Double disk failure is improbable but not impossible. It's not even improbable if the disks are the same kind purchased at the same time.

Yep: if you buy a pair disks together, there's a fair chance they'll both be from the same manufacturing batch, which correlates with disk defects.

This is why I try to mismatch manufacturers in RAID arrays. I'm told there is a small performance hit (things run towards the speed of the slowest, separately in terms of latency and throughput) but I doubt the difference is high and I like the reduction in potential failure-during-rebuild rates. Of course I have off-machine and off-site backups as well as RAID, but having to use them to restore a large array would be a greater inconvenience than just being able to restore the array (followed by checksum verifies over the whole lot for paranoia's sake).

Re: HN is up again

#363
post #307

Earlier quoted context omitted.

You are never going to guess how long the HN SSDs were in the servers... never ever... OK... I'll tell you: 4.5years. I am not even kidding.

It's concerning that a hosting company was unaware of the 40,000 hour situation with SSD it was deploying. Anyone in hosting would have been made aware of this, or at least should have kept a better grip on happenings in the market.

Yeah, this is why you run all equipment in a test environment for 4.5 years before deploying it to prod. Really basic stuff.

Re: HN is up again

#364

Earlier quoted context omitted.

It's concerning that a hosting company was unaware of the 40,000 hour situation with SSD it was deploying. Anyone in hosting would have been made aware of this, or at least should have kept a better grip on happenings in the market.

Yeah, this is why you run all equipment in a test environment for 4.5 years before deploying it to prod. Really basic stuff.

The HD makers started issuing warnings in 2020... this was foreseeable

Re: HN is up again

#365
post #323
post #322

Earlier quoted context omitted.

Wow. It's possible that you have nailed this. Edit: here's why I like this theory. I don't believe that the two disks had similar levels of wear, because the primary server would get more writes than the standby, and we switched between the two so rarely. The idea that they would have failed within hours of each other because of wear doesn't seem plausible. But the two servers were set up at the same time , and it's…

This morning, I googled for issues with the firmware and the model of SSD, I got nothing. But now I am searching for "40000 hours SSD" and a million relevant results. Of course, why would I search for 40000 hours. This thread is making me feel a lot less crazy.

There are times I don't miss dealing with random hardware mystery bullshit.

This one is just ... maddening.

Re: HN is up again

#366
post #353

Earlier quoted context omitted.

This kind of thing is why I love Hacker News. Someone runs into a strange technical situation, and someone else happens to share their own obscure, related anecdote, which just happens to precisely solve the mystery. Really cool to see it benefit HN itself this time.

It's also an example of the dharma of /newest – the rising and falling away of stories that get no attention: HPE releases urgent fix to stop enterprise SSDs conking out at 40K hours - https://news.ycombinator.com/item?id=22706968 - March 2020 (0 comments) HPE SSD flaw will brick hardware after 40k hours - https://news.ycombinator.com/item?id=22697758 - March 2020 (0 comments) Some HP Enterprise SSD will brick after…

Popularity is a very poor relevance / truth heuristic.

Re: HN is up again

#367
post #331
post #309

Earlier quoted context omitted.

Let me narrow my guess: They hit 4 years, 206 days and 16 hours . . . or 40,000 hours. And that they were sold by HP or Dell, and manufactured by SanDisk. Do I win a prize? (None of us win prizes on this one).

I wonder if it might be closer to 40,032 hours. The official Dell wording [1] is "after approximately 40,000 hours of usage". 2^57 nanoseconds is 40031.996687737745 hours. Not sure what's special about 57, but a power of 2 limit for a counter makes sense. That time might include some manufacturer testing too. [1] https://www.reddit.com/r/sysadmin/comments/f5k95v/dell_emc_u...

It might not be nanoseconds, but something that's a power of 2 number of nanoseconds going into an appropriately small container seems likely. For example, a 62.5MHz counter going into 53 bits breaks at the same limit. Why 53 bits? That's where things start to get weird with IEEE doubles - adding 1 no longer fits into the mantissa and the number doesn't change. So maybe someone was doing a bit of fp math to figure out the time or schedule a next event? Anyway, very likely some kind of clock math that wrapped or saturated and broke a fundamental assumption.

Re: HN is up again

#368

Earlier quoted context omitted.

This makes total sense but I've never heard of it. Is there any literature or writing about this phenomenon? I guess proper redundancy is having different brands of equipment also in some cases.

Not sure about literature but that was a known thing in the Ops circles I was in 10 years ago: never use the same brand for disk pairs, to minimize wear-and-tear related defects from arising at the same time.

We used to use the same brand, but different models or at least ensure they were from different manufacturing batches.

Re: HN is up again

#369
post #353

Earlier quoted context omitted.

This kind of thing is why I love Hacker News. Someone runs into a strange technical situation, and someone else happens to share their own obscure, related anecdote, which just happens to precisely solve the mystery. Really cool to see it benefit HN itself this time.

It's also an example of the dharma of /newest – the rising and falling away of stories that get no attention: HPE releases urgent fix to stop enterprise SSDs conking out at 40K hours - https://news.ycombinator.com/item?id=22706968 - March 2020 (0 comments) HPE SSD flaw will brick hardware after 40k hours - https://news.ycombinator.com/item?id=22697758 - March 2020 (0 comments) Some HP Enterprise SSD will brick after…

Easy to imagine why this didn’t capture peoples’ attention in late March 2020…

Re: HN is up again

#370
post #333
post #331

Earlier quoted context omitted.

I wonder if it might be closer to 40,032 hours. The official Dell wording [1] is "after approximately 40,000 hours of usage". 2^57 nanoseconds is 40031.996687737745 hours. Not sure what's special about 57, but a power of 2 limit for a counter makes sense. That time might include some manufacturer testing too. [1] https://www.reddit.com/r/sysadmin/comments/f5k95v/dell_emc_u...

See! People should register via mail for those important notifications! (Or alternatively do quarterly checks that your firmware is up to date).

A lot of companies have teams dedicated to hardware that don’t give a shit about it. And their managers don’t give a shit.

Then the people under them who do give a shit, because they depend on those servers, aren’t allowed to register with HP etc for updates, or to apply firmware updates, because “separation of duties”.

Basically, IT is cancer from the head down.

Post reply on HN