Live data from Hacker News

Nvidia-smi hangs indefinitely after ~66 days

github.com

1–10 of 54 posts

Re: Nvidia-smi hangs indefinitely after ~66 days

#2
a pet peeve of mine, (along with people brigading on issues/threads e.g. posting them to unrelated news sites... op....) is woefully incorrect language.

> at day 66 all our jobs started randomly failing

if there's a definable pattern, you can call it unpredictabily, but you can't call it randomly.

Re: Nvidia-smi hangs indefinitely after ~66 days

#3
Crazy, so if I understand correctly, something with B200s and nvlink is causing issues where after 66 days and 12 hours of uptime, nvidia-smi and other jobs start failing, timing out, then once you restart the cluster it starts working again.

They suspect jobs will work if you only use 1 B200, but one person power cycled so wasn’t able to test it. Hopefully they won’t have to wait another 66 days for further troubleshooting.

Re: Nvidia-smi hangs indefinitely after ~66 days

#4
post #3

Crazy, so if I understand correctly, something with B200s and nvlink is causing issues where after 66 days and 12 hours of uptime, nvidia-smi and other jobs start failing, timing out, then once you restart the cluster it starts working again. They suspect jobs will work if you only use 1 B200, but one person power cycled so wasn’t able to test it. Hopefully they won’t have to wait another 66 days for further troubles…

Some 32-bit counter somewhere used when in NVLINK overflows?

Re: Nvidia-smi hangs indefinitely after ~66 days

#5

a pet peeve of mine, (along with people brigading on issues/threads e.g. posting them to unrelated news sites... op....) is woefully incorrect language. > at day 66 all our jobs started randomly failing if there's a definable pattern, you can call it unpredictabily, but you can't call it randomly.

Unexpectedly is probably what they meant

Re: Nvidia-smi hangs indefinitely after ~66 days

#7
post #3

Crazy, so if I understand correctly, something with B200s and nvlink is causing issues where after 66 days and 12 hours of uptime, nvidia-smi and other jobs start failing, timing out, then once you restart the cluster it starts working again. They suspect jobs will work if you only use 1 B200, but one person power cycled so wasn’t able to test it. Hopefully they won’t have to wait another 66 days for further troubles…

Some 32-bit counter somewhere used when in NVLINK overflows?

Isn't 32bit counter 49 days? Assuming that one was counting milliseconds, at least.

Only remember that because that's the limit for Windows 95…

Post reply on HN