At last year's KubeCon someone presented https://github.com/opsgenie/kubernetes-event-exporter
Tracking OOMKilled counts via the kubelet also had an open issue which was closed without being fixed: https://github.com/kubernetes/kubernetes/pull/87856
21–30 of 33 posts
At last year's KubeCon someone presented https://github.com/opsgenie/kubernetes-event-exporter
Tracking OOMKilled counts via the kubelet also had an open issue which was closed without being fixed: https://github.com/kubernetes/kubernetes/pull/87856
What about healthchecks?
Use serverless guys. k8s is a train wreck of needless complexity for 99% of developers.
Use serverless guys. k8s is a train wreck of needless complexity for 99% of developers.
Another hidden issue is that as a container gets close to running out of memory, it furiously drops read only pages from memory, only to need to read some of them back into memory moments later. This pathological swapping behavior can impact other workloads on the system. cgroups2 has better protections against this behavior.
What about healthchecks?
I don't think it would work to report an OOMkill. The process that would answer the healthcheck probe would be gone.
If the process stops responding to a healthcheck, it's the scheduler's responsibility (k8s in this case) to handle it. Crashes in this case should be handled in a similar way, whether it's due to OOM or a bug.
Maybe I'm missing another scenario here?
Use serverless guys. k8s is a train wreck of needless complexity for 99% of developers.
What about healthchecks?
If a sub-process gets OOMKilled and the container doesn't die, then it's most likely that the parent process didn't handle that scenario. In which case the health-check wouldn't cover that issue.
Another hidden issue is that as a container gets close to running out of memory, it furiously drops read only pages from memory, only to need to read some of them back into memory moments later. This pathological swapping behavior can impact other workloads on the system. cgroups2 has better protections against this behavior.
EarlyOOM [1] is a configurable daemon that kills processes early enough to (hopefully) prevent thrashing. I'm using it on my Linux desktops (it has proven to catch my own programs' runaway memory usage before it risks locking up the development machine), but it may also be useful on servers. It logs to syslog but also can be configured to run a program on kill events.
[1] https://github.com/rfjakob/earlyoom, https://launchpad.net/ubuntu/+source/earlyoom, https://packages.debian.org/search?keywords=earlyoom
(Why would a user space OOM killer be necessary if the kernel has better information about the state of the world? I don't know the details, but my interpretation is that because people disliked OOM killing, the kernel devs made the kernel OOM killer trigger so late that it is largely useless. If that's true and thus a social problem, maybe it needs to be solved on that level, too.)
BTW in my experience, Linux 2.2 used to handle out of memory situations much more gracefully than any later kernel version.