Live data from Hacker News

Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

ubicloud.com

11–20 of 117 posts

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#11
post #5

Earlier quoted context omitted.

Author of the blog post here. Yeah, this is generally a good practice. The silver lining is that our suffering helped uncover the underlying issue faster. :) This isn’t part of the blog post, but we also considered getting the servers and keeping them idle, without actual customer workload, for about a month in the future. This would be more expensive, but it could help identify potential issues without impacting our…

>The silver lining is that our suffering helped uncover the underlying issue faster. Did you actually uncover the true root cause? Or did they finally uncap the power consumption without telling you, just as they neither confirmed nor denied having limited it?

The root cause was a problem with the motherboard, though the exact issue remains unknown to us. I suspect that a component on the motherboard may have been vulnerable to power limitations or fluctuations and that the newer-generation motherboards included additional protection against this. However, this is purely my speculation.

I don't believe they simply lifted a power cap (if there was one in the first place). I genuinely think the fix came after the motherboard replacements. We had 2 batches of motherboard replacements and after that, the issue disappeared.

If someone from Hetzner is here, maybe they can give extra information.

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#12
post #5

Earlier quoted context omitted.

Author of the blog post here. Yeah, this is generally a good practice. The silver lining is that our suffering helped uncover the underlying issue faster. :) This isn’t part of the blog post, but we also considered getting the servers and keeping them idle, without actual customer workload, for about a month in the future. This would be more expensive, but it could help identify potential issues without impacting our…

>The silver lining is that our suffering helped uncover the underlying issue faster. Did you actually uncover the true root cause? Or did they finally uncap the power consumption without telling you, just as they neither confirmed nor denied having limited it?

hetzner is currently replacing motherboards of their dedicated servers [1] But I dont know if thats the same issue that was mentioned in the article.

[1] https://status.hetzner.com/incident/7fae9cca-b38c-4154-8a27-...

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#13
post #5
post #2

> Looking back, waiting six months could have helped us avoid many issues. Early adopters usually find problems that get fixed later. This is really good advice and what I'm following for all systems which need to be stable. If there aren't any security issues, I either wait a few months or keep one or two versions behind.

Author of the blog post here. Yeah, this is generally a good practice. The silver lining is that our suffering helped uncover the underlying issue faster. :) This isn’t part of the blog post, but we also considered getting the servers and keeping them idle, without actual customer workload, for about a month in the future. This would be more expensive, but it could help identify potential issues without impacting our…

Customers are the best QA. And they pay you too, instead of the reverse!

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#14
post #2

> Looking back, waiting six months could have helped us avoid many issues. Early adopters usually find problems that get fixed later. This is really good advice and what I'm following for all systems which need to be stable. If there aren't any security issues, I either wait a few months or keep one or two versions behind.

This is a wildly successfully pattern in nature, the old using the young and inexperienced, as enthusiastic test units. In the wild for example in Forrest, old boars give safety squeaks to send the younglings ahead into a clearing they do not trust. The equivalent to that- would be to write a tech-blog entry that hypes up a technology that is not yet production ready.

Just for curiosity: do you have a source?

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#15
We will never know, but I wonder if it could be a power/signaling or VRM issue - the CPU non getting hot doesn't mean something else on the board has gone out of spec and into catastrophic failure.

Motherboard issues around power/signaling are a pain to diagnose, they will emerge as all sort of problems apparently related to other components (ram failing to initialize and random restarts are very common in my experience) and you end up swapping everything before actually replacing the MB...

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#16
post #12

Earlier quoted context omitted.

>The silver lining is that our suffering helped uncover the underlying issue faster. Did you actually uncover the true root cause? Or did they finally uncap the power consumption without telling you, just as they neither confirmed nor denied having limited it?

hetzner is currently replacing motherboards of their dedicated servers [1] But I dont know if thats the same issue that was mentioned in the article. [1] https://status.hetzner.com/incident/7fae9cca-b38c-4154-8a27-...

Thats the same issue, yes.

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#17
post #4

> To increase the number of machines under power constraints, data center operators usually cap power use per machine. However, this can cause motherboards to degrade more quickly. Can anyone elaborate on this point? This is counter to my intuition (and in fact, what I saw upon a cursory search), which is that power capping should prolong the useful lifetime of various components. The only search results I found that…

Yep, that's weird, I've always read that high power/temp can degrade electronics way faster. Any EE can shed a light here?

As an electronics engineer I have no idea what the author is talking about here and was about to post the same question.

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#18
At a previous company, devops would regularly find CPU fan failures on Hetzner. That's in addition to the usual expected HD/SSD failures. You've got to do your own monitoring, it's one of the reasons why unmanaged servers are cheaper than cloud instances.

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#19
post #10
post #2

> Looking back, waiting six months could have helped us avoid many issues. Early adopters usually find problems that get fixed later. This is really good advice and what I'm following for all systems which need to be stable. If there aren't any security issues, I either wait a few months or keep one or two versions behind.

GitHub is looking to add this feature to dependabot: https://github.com/dependabot/dependabot-core/issues/3651

In theory, that works in practice nope. You get a random update with a possible bug inside that is only fixed by a new version that you won't get until later. The other strategy is to wait for a package to be fully stable (no update), and in that case, some packages that receive daily/weekly updates are never updated

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#20
post #4

> To increase the number of machines under power constraints, data center operators usually cap power use per machine. However, this can cause motherboards to degrade more quickly. Can anyone elaborate on this point? This is counter to my intuition (and in fact, what I saw upon a cursory search), which is that power capping should prolong the useful lifetime of various components. The only search results I found that…

The only place I could find some answer that sheds some light was StackOverflow:

https://electronics.stackexchange.com/a/65827

> A mosfet needs a certain voltage at its gate to turn fully on. 8V is a typical value. A simple driver circuit could get this voltage directly from the power that also feeds the motor. When this voltage is too low to turn the mosfet fully on a dangerous situation (from the point of view of the moseft) can arise: when it is half-on, both the current through it and the voltage across it can be substantial, resulting in a dissipation that can kill it. Death by undervoltage.

Post reply on HN