Live data from Hacker News

Server BMCs can need to be rebooted every so often

utcc.utoronto.ca

11–20 of 96 posts

Re: Server BMCs can need to be rebooted every so often

#12
It appears that a large number of these are rebranded AMI MegaRAC software running on ASPEED processors, which are a little ARM chip with a virtual display card hanging off a PCIe x2 slot embedded into the mainboard.

AFAIK larger vendors like Dell and HP have their own thing.

Re: Server BMCs can need to be rebooted every so often

#13
Experiences with iDRAC 6, 7, 8 have been terrible. After high runtime they stop to respond via HTTP and SNMP, its just rubbish. Reboot of BMC helped Sometimes, sometimes only a powerdrain. Back then even the ProSupport could not do much, they support a system unsupported by their devs.

The latest iDRACs (with the fancy blue GUI) work a little better. I have no numbers to back all that, its just a feeling, maybe the problems are yet to come.

Given the Dell pricing, they should be better. But i've heard from collegues that ILO and others are not much better either.

Re: Server BMCs can need to be rebooted every so often

#14

This is brought to you by the BMC with a KVM-over-IP that wouldn't accept '2' entered on the (virtual) keyboard in any way or form. This is one of those things that I'd probably be willing to spend a ton of time analysing to find the root cause if I could. Dump the memory image, debug the code, and figure out why it's just that (and possibly other) keys. My guess just from the description is that a bitflip happened e…

This stuff seems to happen on proxmox when you try speaking to a vm console using the novnc web interface. If I recall correctly for me it wouldn't handle modifier keys correctly so stuff like shift-2 for @ wouldn't work.

I suspect it is some kind of console and raw keycode mismatch with the remote ui, maybe via the browser.

Re: Server BMCs can need to be rebooted every so often

#15
post #6

There are too many names for the BMCs, even within a single vendor. BMC, ILO, LOM, DRAC, iDRAC, etc And the worst are those that use java applets or webstart and require a ancient java version.

I am honestly surprised how bad many of these are, and in production, no less. I recently set up a supermicro system and spent a whole day just trying to figure out what to install to get the stupid ancient Java crap to load so I could mount an ISO.

We have various Supermicro boards in production at work with BMCs from 2018 or so. The ATEN iKVM on them works just fine with a recent OpenJDK 11 and OpenWebStart. I’ve found that all the features work including mounting ISOs and doing remote upgrades. No need to whine about installing anything ancient or spreading ridiculous Java FUD.

Re: Server BMCs can need to be rebooted every so often

#16
BMCs have to be some of the most unreliable devices that I've worked with. Some of the issues I encountered at my last job:

* [ASRock BMC] The BMC firmware updater sometimes causes the NICs to get stuck in a bad state where every NIC has the same MAC address. This can be resolved via a proprietary UEFI application for reflashing the correct MAC address.

* [Dell iDRAC] Local authentication randomly stops working due to some tmpfs running out of space (can see the message if there's an active SSH session). IPMI/SSH occasionally works well enough to issue a reboot command, but when it doesn't, sending the BMC reset comand to /dev/ipmi0 in the server OS is needed.

* [Dell iDRAC] Setting an asset tag via the IPMI DCMI command has a 15 character limit. If that limit is exceeded, the success response is still returned, but when querying the asset tag, random junk data longer than 15 bytes is returned. If I had to guess, I bet there was an sprintf() call somewhere in there :). This was fixed in newer iDRAC firmware. Now, it stores the last 15 bytes of the asset tag instead of returning an error.

* [Lenovo IMM] The shift or alt key sometimes gets stuck on the emulated keyboard without having used the remote console since the last BMC reboot. Can't be fixed by repressing the button, neither physically nor via the virtual keyboard. BMC reboot required.

* [Lenovo IMM] Booting the UEFI shell sometimes crashes both the system and the BMC.

* [Lenovo IMM] BMC and BIOS update sometimes claims to have succeeded, but didn't actually take effect.

* [Lenovo IMM] Rebooting the BMC via the web UI or SSH sometimes doesn't work. Making 50+ simultaneous requests to the login page is enough to crash and restart some component that allows the BMC reboot command to work again though.

* [Supermicro BMC] The BIOS update sometimes doesn't fully upload, but claims that it did. It still parses the header, so the new/old version fields look correct. Sometimes, rebooting the BMC and reflashing works. Other times, only a USB drive + USB keyboard + a recovery key combination fixes it.

* [Supermicro BMC] The remote console sometimes completely fails to initialize (though I've only seen this on servers where the BMC uptime was measured in years). Not just a blank screen. The GPU device was just... gone.

* [Supermicro BMC] Various IPMI commands just lie about successful execution. For example, setting the asset tag via the FRU succeeds, but has no effect. Those commands require toggling a write lock bit via an OEM command, which I only found by reverse engineering. Other commands, like the set asset tag DCMI command, leave the data in a corrupted state until a BMC reboot.

And finally, not really a bug, but an interesting thing about the Lenovo IMM. Instead of exposing information via standard IPMI features, like FRU or DCMI commands, the Lenovo IMM implements a virtual filesystem over OEM IPMI commands. These are commands, like (my naming) open_ro, open_rw, get_size, read, write, and close. They sometimes fail too. I think I ended up making all commands retry up to 10 times with a 5 second delay. At least Lenovo gets error return values right :).

To query the asset tag, you have to open_ro the "config.efi" file, get the size (because read-until-EOF doesn't always work), do a read loop, and close the file. Then, you have the XML data from the file you can query (20 seconds later due to retries). (If anyone ever needs to deal with the Lenovo IMM programmatically, I'd highly recommend the pyghmi [1] library. Wish I knew about it before reverse engineering their proprietary commands...)

[1] https://opendev.org/x/pyghmi

Re: Server BMCs can need to be rebooted every so often

#17

Earlier quoted context omitted.

I am honestly surprised how bad many of these are, and in production, no less. I recently set up a supermicro system and spent a whole day just trying to figure out what to install to get the stupid ancient Java crap to load so I could mount an ISO.

We have various Supermicro boards in production at work with BMCs from 2018 or so. The ATEN iKVM on them works just fine with a recent OpenJDK 11 and OpenWebStart. I’ve found that all the features work including mounting ISOs and doing remote upgrades. No need to whine about installing anything ancient or spreading ridiculous Java FUD.

With current browsers Java applets are not supported anymore. Some older HPE systems the didn't update the firmware to provide alternatives.

I've seen multiple vendors with problematic code that didn't work with newer Java versions.

This is by no means meant to bash Java. Some non-Java BMCs can be horrible as well (e.g. require many TCP ports in a firewall/tunnel unfriendly way or require SSH with old algorithms that are no longer enabled by default, or telnet..)

Re: Server BMCs can need to be rebooted every so often

#18
A few of my recent systems have come with a built-in combo BMC on the motherboard NICs. I haven't seen this before, is it a new trend? I'm imagining they put in a switch in there with the NICs? How does this even work?

I'm doing a few timing sensitive projects involving hardware timestamping in the NICs. Does this mess with say timing variability? I've disabled the BMCs out of paranoia but I don't know the topology inside.

Re: Server BMCs can need to be rebooted every so often

#19
post #18

A few of my recent systems have come with a built-in combo BMC on the motherboard NICs. I haven't seen this before, is it a new trend? I'm imagining they put in a switch in there with the NICs? How does this even work? I'm doing a few timing sensitive projects involving hardware timestamping in the NICs. Does this mess with say timing variability? I've disabled the BMCs out of paranoia but I don't know the topology i…

NC-SI. It allows the BMC and the host to share a single physical network connection.

Re: Server BMCs can need to be rebooted every so often

#20
A colleague from a different department was managing a fleet of servers doing a lot of computing. Their BMCs would just stop working after a few weeks of uptime, every time. He was annoyed by it but mostly ignored it as it was mostly just used for additional monitoring. The power outlets in the racks were remote controllable, so you could still hard reset a server if this was required.

However, as this fleet was used for computation, after a while they noticed that whenever the BMC stopped working, performance of the system increased by almost 10% or so. Definitely a non-neglible amount. So they kept the machines running in the broken BMC state for as much as possible.

Post reply on HN