Live data from Hacker News

Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

ubicloud.com

91–100 of 117 posts

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#91
post #74

Earlier quoted context omitted.

However, eco-friendly power modes can reduce electricity usage, so they can be friendlier for our climate. https://www.rvo.nl/onderwerpen/energie-besparen-de-industrie...

I'm not sure why you are downvoted. Is it wrong?

Probably blowback from "environmentalism at any cost" thinking.

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#92
post #33
post #28

> One of the providers we like is Hetzner because of their affordable and reliable servers. > In the days that followed, the crash frequency increased. I don't find the article conclusive whether they would still call them reliable.

To their credit they actually fixed the problem. Good luck getting this level of support from any of the big 3 public cloud providers.

The main difference being that you talk with real humans who try to help you, not computer programs designed to give you an illusion…

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#93

At a previous company, devops would regularly find CPU fan failures on Hetzner. That's in addition to the usual expected HD/SSD failures. You've got to do your own monitoring, it's one of the reasons why unmanaged servers are cheaper than cloud instances.

No? Maybe you cloud kids don't know how this stuff works, but unmanaged just means you get silicon-level access and remote KVM. It's still the hosting company's responsibility to competently own, maintain, and repair the physical hardware. That includes monitoring. In the old days you had to run a script or install a package to hook into their monitoring....but with IPMI et al being standard they don't need anything…

Alright, let's say the hosting company has an out-of-band mechanism for detecting reboots. How do they know if the reboots are abnormal (like in this case) or normal, customer-ordered reboots after software upgrades?

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#94

Earlier quoted context omitted.

The term PROCHOT just brought me back to vivid memories of debugging exactly that at Facebook a while ago. It was very non-obvious to debug since pretty much most emitted metrics, apart from mysterious errors/timeouts to our service, looked reasonable. Even the cpu usage and cpu temperature graphs looked normal since it was a bogus prochot and not actually a real thermal throttling

And it brought me back to memories of debugging that on my friends laptop. It kept going to 400mhz.. i suspected throttling and we got it cleaned thermal paste replaced and all that. Still throttled. We replaced the windows with linux since it was atleast a bit more usable At the time I didn't know about PROCHOT. And my googling skills clearly weren't sufficient. One fine day during lunch at a place on campus, Id rea…

I once had a dell laptop that after about three years started complaining on power up that I wouldn't be using a genuine dell PSU and should switch to one. I ignored it at first because you could just hit enter and carry on, but after a while I noticed that every time this happened, the cpu would clock at a fixed 800mhz. I ordered a new power brick but the message didn't go away, so I returned the brick and decided to never buy dell again.

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#95
post #69

Dell has this problem sometimes. I remember getting the first batch one of their older servers when they were new. We had to replace motherboards' I/O (rear) section because the servers lost some devices on that part (e.g.: Ethernet controllers, iDRAC, sometimes BIOS) for some time. After shaking out these problems, they ran for almost a decade. We recently retired them because we worn down everything on these server…

Dell has tons of issues. A faulty mini board of the front led can actually stop the server from booting/running at all (even drac will be dead)

Interesting. From my experience, Dell is generally one of the least problematic brands when compared in large numbers.

Another surprising name is Huawei. Their servers just don't die.

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#96
post #93

Earlier quoted context omitted.

No? Maybe you cloud kids don't know how this stuff works, but unmanaged just means you get silicon-level access and remote KVM. It's still the hosting company's responsibility to competently own, maintain, and repair the physical hardware. That includes monitoring. In the old days you had to run a script or install a package to hook into their monitoring....but with IPMI et al being standard they don't need anything…

Alright, let's say the hosting company has an out-of-band mechanism for detecting reboots. How do they know if the reboots are abnormal (like in this case) or normal, customer-ordered reboots after software upgrades?

Probably covered here:

> In the old days you had to run a script or install a package to hook into their monitoring....but with IPMI et al being standard they don't need anything from you to do their job

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#97
post #93

Earlier quoted context omitted.

Alright, let's say the hosting company has an out-of-band mechanism for detecting reboots. How do they know if the reboots are abnormal (like in this case) or normal, customer-ordered reboots after software upgrades?

Probably covered here: > In the old days you had to run a script or install a package to hook into their monitoring....but with IPMI et al being standard they don't need anything from you to do their job

How can IPMI detect the cause (kernel panic vs user command) for restart?

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#98
post #28

> One of the providers we like is Hetzner because of their affordable and reliable servers. > In the days that followed, the crash frequency increased. I don't find the article conclusive whether they would still call them reliable.

Hetzner's reliable... until they aren't. Since they don't do any sort of monitoring on their bare metal servers at all, at least insofar as I can tell having been a customer of theirs for ten years, you don't know there's as problem until there's a problem, or unless you've got your own monitoring solution in place.

Back in 2012 it regularly happened that we called them because the network was gone, because our monitoring seemed to be better. Or at least quicker than what they showed.

Back in 2006 my coworker claimed he was the person responsible for them adding a "exchange my dead HDD" menu point on the support site because he wrote one of those tickets per week.

When I got a physical server, the HDD died in the first 48h, so I've not exactly forgiven them, even if this was a tragic story over the last 18 or so years...

On the other hand, I've been recommending their cloud vps for a couple of years because unlike with their HW, I've never had problems.

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#99
post #46

Similar thing happened to a AX102 I currently use, something related the network card which caused crashes. Thankfully hetzner support was helpful with replacement hardware. caused quite some grief but at least it was a good lesson in hardware troubleshooting. Worth it to me personally

Yep same here. AX102 crashes with almost no load, nothing in the logs, won't come on. Hetzner looked at it multiple times and found either nothing or replaced cpu paste or a PSU connector. I migrated to AX162 and so far so good

Same here. Hetzner found no issues with hardware in diagnostics, they insisted it is related to OS/Software side but on my request they changed hardware which fixed issue.

Re: Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode

#100

i am so glad my sign up process with hetzner failed when i was so dumb that i wanted to give them a chance even with the internet full of horrific stories of bad experiences from their customers. lucky me.

Hetzner is fine for what it is, you just need to know that it's all on you and only YOU . YOU do the monitoring. YOU do the troubleshooting. YOU etc., etc. If that doesn't appeal to you, or if you don't have the requisite knowledge, which I admit is fairly broad and encompassing, then it's not for you. For those of you that meet those checkboxes, they're a pretty amazing deal. Where else could I get a 4c/8t CPU with…

We had been colocating servers from decades but there is too much "YOU", compared to that we find Hetzner doing a lot for us (hardware inventory, replacement, remote hands, networking etc). We are slowly moving away from colocating to renting at Hetzner. It is so much better.
Post reply on HN