Live data from Hacker News

EPYC 7002 CPUs may hang after 1042 days of uptime

old.reddit.com

31–40 of 109 posts

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#31
post #26

I feel like some of the comments here are missing the point. Yes it's only likely to effect a small number of users, so did the intel fdiv bug, both are defective products. Back then intel were pressured into a recall, today we seem too willing to put up with being sold broken stuff.

> today we seem too willing to put up with being sold broken stuff. i remember reading that when hard disks just came into the mass market they were so expensive that having some bad sectors was not such a big deal... and so hard disk would usually come with a sheet of paper listing the known broken sectors (detected at QA stage, i guess). maybe someone older than me (i guess somebody in their 50ies or 60ies) could c…

I'm not that old, but I remember seeing bad sector lists as stickers on some hard disks.

I'm not sure if that ever went away, though... I think the IDE firmware in more modern hard disks knew how to redirect bad sectors to good sectors, so the end user never even noticed.

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#32

Earlier quoted context omitted.

It feels a little bit different. One creates uncertainty in all floating point results, given you don’t know when it happens. The other requires you to reboot maybe every ~3 years and you know exactly when it happens. I’m not saying we should tolerate a defect, but it doesn’t feel nearly as problematic.

It means anyone launching amd powered virtual machines on cloud providers can experience this now, at any point, and you don't know when it will happen, given this type of CPU could have been bought, booted or rebooted anytime in the past three years. Seems comparably problematic to me.

I’d prefer a bug that crashes a program than one that quietly inserts wrong data and keeps going.

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#33

I feel like some of the comments here are missing the point. Yes it's only likely to effect a small number of users, so did the intel fdiv bug, both are defective products. Back then intel were pressured into a recall, today we seem too willing to put up with being sold broken stuff.

There are always bugs in silicon, just like there are bugs in software. They mostly show up under "a highly specific and detailed set of internal timing conditions". There are 40 documented erratas on EPYC 7002s alone; there are 35 in the 13gen Intel CPUs, including, curiously, RPL038, "Processor Exiting Package C6 or C8 May Hang". Mobile ARM chip manufacturers are notoriously bad at documenting their bugs, so who knows how many they have.

This one is interesting because its preconditions are so trivial, and it will affect many more people than usual.

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#34

I feel like some of the comments here are missing the point. Yes it's only likely to effect a small number of users, so did the intel fdiv bug, both are defective products. Back then intel were pressured into a recall, today we seem too willing to put up with being sold broken stuff.

It feels a little bit different. One creates uncertainty in all floating point results, given you don’t know when it happens. The other requires you to reboot maybe every ~3 years and you know exactly when it happens. I’m not saying we should tolerate a defect, but it doesn’t feel nearly as problematic.

It also has a fairly easy solution: disable the CC6 sleep state. The practical effects from that will most likely be minimal or non-existent for most users of these CPUs.

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#35

Earlier quoted context omitted.

It feels a little bit different. One creates uncertainty in all floating point results, given you don’t know when it happens. The other requires you to reboot maybe every ~3 years and you know exactly when it happens. I’m not saying we should tolerate a defect, but it doesn’t feel nearly as problematic.

It means anyone launching amd powered virtual machines on cloud providers can experience this now, at any point, and you don't know when it will happen, given this type of CPU could have been bought, booted or rebooted anytime in the past three years. Seems comparably problematic to me.

The cloud providers now know of this bug. They will live migrate you to a different machine or shutdown reboot. Only on-premise will have this issue.

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#37
post #13
post #3

Earlier quoted context omitted.

Air gapped machines and kernel live patching both exist.

And how many people use that? Most servers today are not air-gapped.

I have kernel live patching on my mother's computer because it means she has to know how to do less

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#38

Earlier quoted context omitted.

It means anyone launching amd powered virtual machines on cloud providers can experience this now, at any point, and you don't know when it will happen, given this type of CPU could have been bought, booted or rebooted anytime in the past three years. Seems comparably problematic to me.

I’d prefer a bug that crashes a program than one that quietly inserts wrong data and keeps going.

It probably depends on your workload, which is a bigger deal. The fdiv bug was pretty bad, but at least fixable in software (at some cost). Anyway, recall is the right decision in either case (unless there’s a good enough workaround).

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#39

Earlier quoted context omitted.

You live in your own World with other people. Please just keep in mind there are many other Worlds with other people and laws of the Universe. I don't know if you're young or don't know much about history but what you describe is a fairly recent way of looking at things, it's not the only one and I guarantee you it will become "out of fashion".

Yeah, the "cattle not pets" philosophy is fairly recent, but I don't see it changing any time soon. If anything we're going even more in that direction. And it makes a lot of sense because if uptime is that important, then no matter how fancy the hardware it can't do anything about disasters or losing internet connectivity.

We might go so far in that direction we wind up right back on the other side. It always happens, it’s more of a pendulum swinging back and forth than any kind of straight forward progress you are imagining.

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#40

Earlier quoted context omitted.

It feels a little bit different. One creates uncertainty in all floating point results, given you don’t know when it happens. The other requires you to reboot maybe every ~3 years and you know exactly when it happens. I’m not saying we should tolerate a defect, but it doesn’t feel nearly as problematic.

It means anyone launching amd powered virtual machines on cloud providers can experience this now, at any point, and you don't know when it will happen, given this type of CPU could have been bought, booted or rebooted anytime in the past three years. Seems comparably problematic to me.

> amd powered virtual machines on cloud providers

Cloud providers are very unlikely to use sleep states. I mean, is possible... but I'd bet against it.

Post reply on HN