Live data from Hacker News

EPYC 7002 CPUs may hang after 1042 days of uptime

old.reddit.com

21–30 of 109 posts

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#21

Earlier quoted context omitted.

I dont understand opinons like this Just because it would be dangerous for your nodejs web_app.exe running on ubuntu behind apache fully exposed on the internet then there are billion other ways to use computers, like even air gapped systems. So, dont try to justify obvious flaw

I mean, hardware is cheap enough that any server of importance should be individually disposable. Yeah, you can do stuff to maximize uptime but if it needs to stay up that badly you have to consider the case of the hardware needing to be turned off at some point. > So, dont try to justify obvious flaw I'm not, it's a bug and should be fixed. But I think if anything is powered for 3 years straight it's a bit concernin…

You live in your own World with other people. Please just keep in mind there are many other Worlds with other people and laws of the Universe.

I don't know if you're young or don't know much about history but what you describe is a fairly recent way of looking at things, it's not the only one and I guarantee you it will become "out of fashion".

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#22

I feel like some of the comments here are missing the point. Yes it's only likely to effect a small number of users, so did the intel fdiv bug, both are defective products. Back then intel were pressured into a recall, today we seem too willing to put up with being sold broken stuff.

It feels a little bit different.

One creates uncertainty in all floating point results, given you don’t know when it happens. The other requires you to reboot maybe every ~3 years and you know exactly when it happens.

I’m not saying we should tolerate a defect, but it doesn’t feel nearly as problematic.

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#23
post #8
post #5

It seems a C6 state is an individual core sleeping. The intersection of people who don't reboot for 3 years and people who have sleep states enabled must be pretty small. It's an interesting bug though!

I had a very similar issue with some AMD-based servers (bulldozer, I think) about ten years ago. There was a bug where Xen-based virtual machines could set a C-state on cores it was assigned, but for whatever reason it wasn't able to wake them up. It was fun trying to figure out what the heck was going on.

I have C-states already disabled because of old linux kernel bug where the kernel hang on Zen3 architecture. So not much to see here :)

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#25

Earlier quoted context omitted.

I mean, hardware is cheap enough that any server of importance should be individually disposable. Yeah, you can do stuff to maximize uptime but if it needs to stay up that badly you have to consider the case of the hardware needing to be turned off at some point. > So, dont try to justify obvious flaw I'm not, it's a bug and should be fixed. But I think if anything is powered for 3 years straight it's a bit concernin…

You live in your own World with other people. Please just keep in mind there are many other Worlds with other people and laws of the Universe. I don't know if you're young or don't know much about history but what you describe is a fairly recent way of looking at things, it's not the only one and I guarantee you it will become "out of fashion".

Yeah, the "cattle not pets" philosophy is fairly recent, but I don't see it changing any time soon. If anything we're going even more in that direction.

And it makes a lot of sense because if uptime is that important, then no matter how fancy the hardware it can't do anything about disasters or losing internet connectivity.

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#26

I feel like some of the comments here are missing the point. Yes it's only likely to effect a small number of users, so did the intel fdiv bug, both are defective products. Back then intel were pressured into a recall, today we seem too willing to put up with being sold broken stuff.

> today we seem too willing to put up with being sold broken stuff.

i remember reading that when hard disks just came into the mass market they were so expensive that having some bad sectors was not such a big deal... and so hard disk would usually come with a sheet of paper listing the known broken sectors (detected at QA stage, i guess).

maybe someone older than me (i guess somebody in their 50ies or 60ies) could confirm that.

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#27
post #11

Earlier quoted context omitted.

They're not that rare. Also, there are a lot of other updates that in practice should be followed up with a reboot. For example, any library consumed by systemd (such as openssl) usually requires pid1 to relaunch. For example, debian released an openssl update just yesterday. You can run "checkrestart -v" to try to figure out how to restart every affected app but you'll quickly run into systemd's init process running…

> For example, any library consumed by systemd (such as openssl) usually requires pid1 to relaunch. That does not require a reboot, `systemctl daemon-reexec` is enough.

nice username, can i ask you what did you see?

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#28

I feel like some of the comments here are missing the point. Yes it's only likely to effect a small number of users, so did the intel fdiv bug, both are defective products. Back then intel were pressured into a recall, today we seem too willing to put up with being sold broken stuff.

It feels a little bit different. One creates uncertainty in all floating point results, given you don’t know when it happens. The other requires you to reboot maybe every ~3 years and you know exactly when it happens. I’m not saying we should tolerate a defect, but it doesn’t feel nearly as problematic.

It means anyone launching amd powered virtual machines on cloud providers can experience this now, at any point, and you don't know when it will happen, given this type of CPU could have been bought, booted or rebooted anytime in the past three years.

Seems comparably problematic to me.

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#29
post #13
post #3

Earlier quoted context omitted.

Air gapped machines and kernel live patching both exist.

And how many people use that? Most servers today are not air-gapped.

How many examples will you need before you say "oh ok, I can see some valid concerns."?

I've worked in places where expensive Lab equipment is running off outdated PCs/servers because updates aren't available and they will absolutely stay on for as long as possible.

We're not all silicon valley, things can be expensive and difficult to replace...

Re: EPYC 7002 CPUs may hang after 1042 days of uptime

#30

I feel like some of the comments here are missing the point. Yes it's only likely to effect a small number of users, so did the intel fdiv bug, both are defective products. Back then intel were pressured into a recall, today we seem too willing to put up with being sold broken stuff.

It feels a little bit different. One creates uncertainty in all floating point results, given you don’t know when it happens. The other requires you to reboot maybe every ~3 years and you know exactly when it happens. I’m not saying we should tolerate a defect, but it doesn’t feel nearly as problematic.

Especially compared to something like this: https://www.theregister.com/2020/04/02/boeing_787_power_cycl...
Post reply on HN