Live data from Hacker News

Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

ubicloud.com

101–110 of 117 posts

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#101
post #39

Most other AX models (AX42, AX52 and AX102) also have serious reliability issues, where they will fail after some months. They are based on a faulty motherboard. Hetzner has to replace most, if not all, motherboards for servers built before a certain date over the next 12 months [0] [0] https://docs.hetzner.com/robot/dedicated-server/general-info...

I have two AX42's. One has been stable since I got it during the Eurocup discount period. The other got replaced 2 times so far, but it looks like the latest replacement is holding up. So, it's like 50% failure rate based on my small sample. I guess only Hetzner and ASRock know the real numbers.

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#102

At a previous company, devops would regularly find CPU fan failures on Hetzner. That's in addition to the usual expected HD/SSD failures. You've got to do your own monitoring, it's one of the reasons why unmanaged servers are cheaper than cloud instances.

No? Maybe you cloud kids don't know how this stuff works, but unmanaged just means you get silicon-level access and remote KVM. It's still the hosting company's responsibility to competently own, maintain, and repair the physical hardware. That includes monitoring. In the old days you had to run a script or install a package to hook into their monitoring....but with IPMI et al being standard they don't need anything…

Do Hetzner servers even run IPMI?

For dedicated servers, you have to schedule KVM access in advance, so I assume they need to move some hardware and plug into to your server.

This would mean that IPMI is most likely not available or disabled.

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#103
post #82
post #5

Earlier quoted context omitted.

Author of the blog post here. Yeah, this is generally a good practice. The silver lining is that our suffering helped uncover the underlying issue faster. :) This isn’t part of the blog post, but we also considered getting the servers and keeping them idle, without actual customer workload, for about a month in the future. This would be more expensive, but it could help identify potential issues without impacting our…

Were you able to identify the manufacturer and model/revision of the failing motherboards? This would be extremely helpful when shopping for seconds hand servers.

I cannot find the link now, but it was mentioned that it was ASRock mobos.

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#104

Would anybody with data center experience be able to hazard a guess on what type of commercial resolution Hetzner would have reached with the Motherboard supplier here? Would we assume all mobos replaced free of charge plus compensation?

I think they probably got a batch of these really cheap in the first place, because those servers were offered without the setup fee initially. It was during the soccer World Cup in Germany.

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#105

Earlier quoted context omitted.

No? Maybe you cloud kids don't know how this stuff works, but unmanaged just means you get silicon-level access and remote KVM. It's still the hosting company's responsibility to competently own, maintain, and repair the physical hardware. That includes monitoring. In the old days you had to run a script or install a package to hook into their monitoring....but with IPMI et al being standard they don't need anything…

Do Hetzner servers even run IPMI? For dedicated servers, you have to schedule KVM access in advance, so I assume they need to move some hardware and plug into to your server. This would mean that IPMI is most likely not available or disabled.

Not anymore, but you can abuse pstore to know about last messages from before reboot

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#106

Earlier quoted context omitted.

Is there any downside to ondemand? If your servers aren't running at 100% then there's no point wasting watts, even if you aren't paying for them, right?

It performs terrible if you have an intermittent workload. Like e.g. a request response workload where request processing is cheap (so that the time to increase the frequency matters). I've seen cases it's a more than 2x request latency increase. It can be pretty annoying, because it means that systems can perform better under higher load and that you get drastically different latency depending on whether a request i…

Intermittent workloads on the order of 2 milliseconds, you mean? Frequency scaling is much faster than blinking, as well as ubiquitous in all recent consumer hardware that I'm aware of due to the huge amount of power it saves. Turning it off would, to me, only make sense if you want a server to handle thousands of fast requests per second, but those requests don't come in for periods of, say, 50 ms at a time and so the CPU scales back.

Actually, thinking this through, even then it doesn't make much sense to me: if you have that many short requests coming in, the CPU would simply never scale back if it's reasonably constant. It would first need to see some gap, and why not scale the CPU back in that gap (at the cost of having the 1st request of the next batch be a few milliseconds slower)? From there on, every subsequent request is fast again until there's another lull. Keeping the CPU always on high frequency should only be needed if you have a very tight deadline on that surprise request (high-frequency trading perhaps?), or if your requests are coincidentally always spaced by the same amount of time as CPU scaling measures across. I'm sure these things exist but "intermittent workload" is 90% of all workloads and most workloads definitely aren't meaningfully impacted by cpu scaling

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#107
post #69

Earlier quoted context omitted.

Dell has tons of issues. A faulty mini board of the front led can actually stop the server from booting/running at all (even drac will be dead)

Interesting. From my experience, Dell is generally one of the least problematic brands when compared in large numbers. Another surprising name is Huawei. Their servers just don't die.

well tbf their server pro support is actually good, but we still had a lot of minor issues, we barely had a dead one in Production. Most problems arises after unpacking. Like the one I told. Of course we had dead transrecievers and dead hard drives, but the pro support guys only want the support zip that you can create from the drac and they ask for some details and than they order replacements parts for you.

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#108
post #107

Earlier quoted context omitted.

Interesting. From my experience, Dell is generally one of the least problematic brands when compared in large numbers. Another surprising name is Huawei. Their servers just don't die.

well tbf their server pro support is actually good, but we still had a lot of minor issues, we barely had a dead one in Production. Most problems arises after unpacking. Like the one I told. Of course we had dead transrecievers and dead hard drives, but the pro support guys only want the support zip that you can create from the drac and they ask for some details and than they order replacements parts for you.

Ah, I understand what you go through now. In our case, they came with the parts that we said we gonna need, see that whatever device is dead with their eyes, and just replaced the problematic part.

When the BIOS or the iDRAC is shot, there's no way they gonna get their support ZIP file. If they want they can connect that dead I/O board to a spare part and try. :)

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#109
post #91
post #74

Earlier quoted context omitted.

I'm not sure why you are downvoted. Is it wrong?

Probably blowback from "environmentalism at any cost" thinking.

I’m not downvoted at the moment, but in any case, it’s not really ‘at any cost’. Salient points from the linked report [1]:

* Data centers in the Netherlands use approx. 2% of nationwide electricity production (4% in the US [2])

* Data center electricity usage is nearly constant, while access patterns aren’t

* Even heavily used servers spend 1/3 of power usage on idle cycles, 99% for the most lightly used servers

* Power-saving modes save approx. 10% of electricity without affecting application performance

* Many respondents do not use power-saving modes because of a lack of knowledge, because they fear the consequences, or because they have been instructed not to by their sysadmin/vendor

* Nonetheless, latency-sensitive applications (e.g. HPC or HFT) are not well-suited to power-saving modes

Given these results, it seems sensible to use power-saving modes by default, unless your workload is extremely latency-sensitive.

In any case, I disagree that potential 10% electricity savings across the worldwide data center industry, without affecting application performance, are ‘environmentalism at any cost’.

[1, Dutch] Harryvan, D. et al. (2020). Analyse LEAP Track 1 “Powermanagement.” Rijksdienst voor Ondernemend Nederland. URL: https://www.rvo.nl/sites/default/files/2021/01/Rapport%20LEA...

[1, English] Harryvan, D. et al. (2021). Analysis LEAP Track 1 “Powermanagement.” Netherlands Enterprise Agency. URL: https://www.rvo.nl/sites/default/files/2021/01/Rapport%20LEA...

[2] Shehabi, A. et al. (2024) United States Data Center Energy Usage Report (page 5). Berkeley Lab. URL: https://eta-publications.lbl.gov/sites/default/files/2024-12...

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#110

Earlier quoted context omitted.

Related and perhaps useful: I’ve seen this in multiple cloud offerings already, where the cpu scaling governor is set to some eco-friendly value, in benefit to the cloud provider and in zero benefit to you and much reduced peak cpu perf. To check, run `cat /sys/devices/system/cpu/cpu /cpufreq/scaling_governor`. It should be `performance`. If it’s not, set it with `echo performance | sudo tee /sys/devices/system/cpu/c…

You can tune the ondemand (or any other) governor first to ramp up faster and clock down slower. "performance" should be seen as the nuclear option.

How? I've been wondering the same thing for my laptop. My laptop doesn't even seem to support the ondemand option. It just switches from performance to powersave when I unplug the charger, and becomes painfully slow.
Post reply on HN