In my current workplace (BigCo), we know exactly what's wrong with our alert system. We get alerts that we can't shut off, because they (legitimately) represent customer downtime, and whose root cause we either can't identify (lack of observability infrastructure) or can't fix (the fix is non-trivial and management won't prioritize). Running on-call well is a culture problem. You need management to prioritize observa…
I completely agree that technical tools cannot fix culture problems. However, one of the things that I noticed in my previous companies was that my management chain wasn't even aware that the problem was this bad. We also wanted to add better reporting (like the alert analytics) so that people have more visibility into the state of alerts + on-call load on engineers. What strategies have worked well for you when it c…
"We're paying down our technical debt"