Live data from Hacker News

EPYC 7002 CPUs may hang after 1042 days of uptime

old.reddit.com

91–100 of 109 posts

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#91

Earlier quoted context omitted.

I dont understand opinons like this Just because it would be dangerous for your nodejs web_app.exe running on ubuntu behind apache fully exposed on the internet then there are billion other ways to use computers, like even air gapped systems. So, dont try to justify obvious flaw

I mean, hardware is cheap enough that any server of importance should be individually disposable. Yeah, you can do stuff to maximize uptime but if it needs to stay up that badly you have to consider the case of the hardware needing to be turned off at some point. > So, dont try to justify obvious flaw I'm not, it's a bug and should be fixed. But I think if anything is powered for 3 years straight it's a bit concernin…

Individually disposable, yes. But if you have a cluster of those, and you powered them on at the same time -- as it often happens -- you're in for an exciting ride when your servers start rebooting almost simultaneously, give or take a few minutes.

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#92

Earlier quoted context omitted.

It feels a little bit different. One creates uncertainty in all floating point results, given you don’t know when it happens. The other requires you to reboot maybe every ~3 years and you know exactly when it happens. I’m not saying we should tolerate a defect, but it doesn’t feel nearly as problematic.

It means anyone launching amd powered virtual machines on cloud providers can experience this now, at any point, and you don't know when it will happen, given this type of CPU could have been bought, booted or rebooted anytime in the past three years. Seems comparably problematic to me.

[deleted]

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#93
post #74

Earlier quoted context omitted.

> Back then intel were pressured into a recall, today we seem too willing to put up with being sold broken stuff. This bug only applies to servers that haven’t been rebooted for 3 years and have the CC6 sleep state enabled. It can be worked around by disabling CC6 sleep state or rebooting once every 3 years. If you think operators of these servers can’t be bothered to update and reboot their machines once in 3 years…

I’m picturing a long 50’ aisle filled with racks and a guy with a huge box marked “replacement CPUs” and a screwdriver. Good lord, can you imagine how long just a few of those would take in a data center?

I remember coming to work one morning and having staff at two tables with boxes of RSA keys, and swapping everyones...

(they replaced 40 million of those things..)

https://arstechnica.com/information-technology/2011/06/rsa-f...

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#94
post #34

Earlier quoted context omitted.

It feels a little bit different. One creates uncertainty in all floating point results, given you don’t know when it happens. The other requires you to reboot maybe every ~3 years and you know exactly when it happens. I’m not saying we should tolerate a defect, but it doesn’t feel nearly as problematic.

It also has a fairly easy solution: disable the CC6 sleep state. The practical effects from that will most likely be minimal or non-existent for most users of these CPUs.

I can already see the pain of the myriads of compliance (to all energy reduction directives, at least in EU) people getting strangely obtuse notes from their sw/hw/platform teams, saying in essence, errrrr we need to amend our already thick justification folder, to disable a specific sleep state. I feel a migraine (or a kind of sketch) coming. 'oh and BTW we're field upgrading the whole fleet'.

I guess fighting tooth and nail to disable any and all of these sleep states from the get go is worth it...

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#96
post #34

Earlier quoted context omitted.

It also has a fairly easy solution: disable the CC6 sleep state. The practical effects from that will most likely be minimal or non-existent for most users of these CPUs.

I can already see the pain of the myriads of compliance (to all energy reduction directives, at least in EU) people getting strangely obtuse notes from their sw/hw/platform teams, saying in essence, errrrr we need to amend our already thick justification folder, to disable a specific sleep state. I feel a migraine (or a kind of sketch) coming. 'oh and BTW we're field upgrading the whole fleet'. I guess fighting tooth…

Would this qualify as more CPU errata?

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#97

Earlier quoted context omitted.

Sure, was looking for electrical devices, a better example of what great engineering can achieve I suppose it's the Pons Fabricius [1], bridge built 2,085 years ago, still in use. The problem is, if we can't expect software to run essentially forever, to update without 'restarts', and so forth, how are we ever going to achieve neural chip implants, artificial organs, synthetic agents mining ore in outer space, and so…

Don't get me wrong, I'm not saying it's not a problem. It should be fixed. But if you need a single system to stay up for 3 years straight that's probably not good. There's too much going on in a modern high tech server for that to be a good idea. Everything has a CPU in it (including disks, video cards, network cards, etc). And any of that could make your system unusable by hitting some rare condition. > The problem…

"Any plan where there's a crucial component that must not stop even for a second isn't a very good plan."

Our bodies, just think of our hearts or lungs, don't stop for even a second for 80 something years, and even that 80 is most probably arbitrary with very few changes in cellular control (instead of cancer, cooperate; instead of scar, regenerate [1]). No current software artifact can boast with such a performance. That's the main issue, our technology does not establish a hierarchy of competence [2], where each layer is independently able to solve problems such as the cell-tissue-organ-organism continuum. We must start digitizing the material, assemble assemblers that can assemble themselves [3].

[1] Dr. Michael Levin: Xenobots, Limb Regeneration, and The Power of Cellular Communication, https://www.youtube.com/watch?v=H_TyON2xWeQ

[2] Michael Levin, What do bodies think about?, https://www.youtube.com/watch?v=CVr1OkDqnmo "Nested Cognition, not Merely Structure" starts at 4:32

[3] Neil Gershenfeld, How to Make Almost Anything, The Digital Fabrication Revolution, http://cba.mit.edu/docs/papers/12.09.FA.pdf

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#98

Earlier quoted context omitted.

Don't get me wrong, I'm not saying it's not a problem. It should be fixed. But if you need a single system to stay up for 3 years straight that's probably not good. There's too much going on in a modern high tech server for that to be a good idea. Everything has a CPU in it (including disks, video cards, network cards, etc). And any of that could make your system unusable by hitting some rare condition. > The problem…

"Any plan where there's a crucial component that must not stop even for a second isn't a very good plan." Our bodies, just think of our hearts or lungs, don't stop for even a second for 80 something years, and even that 80 is most probably arbitrary with very few changes in cellular control (instead of cancer, cooperate; instead of scar, regenerate [1]). No current software artifact can boast with such a performance.…

Our bodies actually have a good amount of redundancy.

The cardiac pacemaker (as in the tissue that sets the heart rate) is redundant. There's a primary and a secondary, and both are made of many cells which can take some damage and the entire system will still work.

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#99

Earlier quoted context omitted.

The 7002 seems like it could be used in a workstation, where the “cattle vs pets” thing is less of a distinction, right? (I guess a workstation is sort of like a work dog in this analogy).

That's the EPYC lineup, which is the server model. Support for terabytes of RAM, 128 PCIe lanes, that sort of thing. I mean you could use it in a workstation, but unless you need 4 video cards locally it's probably overkill for most uses. And a workstation should have no problem rebooting once in a while.

You have the whole "I don't understand why something is this way, therefore everyone who does understand why it's that way is wrong" stick down cold. It's not a good look.

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#100
post #70

A machine staying up for almost 3 years is irresponsible in this day and age. Yeah, I remember people having uptime competitions on Slashdot and the like some decades back, but you only need to look at the ssh logs of a 5 minutes old machine to realize this is a terrible idea in modern times.

> A machine staying up for almost 3 years is irresponsible in this day and age. [...] but you only need to look at the ssh logs of a 5 minutes old machine to realize this is a terrible idea in modern times. You don't need to reboot a machine to update ssh. You only need to reboot the machine to update the kernel; for everything else, you just have to restart the corresponding user-space processes (and even PID1 can r…

As I recall, machines made by Tandem Computers, among other highly fault tolerant machines that have regrettably fallen out of fashion, didn't have to reboot even to replace the kernel. They didn't run Linux, tho.
Post reply on HN