In my current workplace (BigCo), we know exactly what's wrong with our alert system. We get alerts that we can't shut off, because they (legitimately) represent customer downtime, and whose root cause we either can't identify (lack of observability infrastructure) or can't fix (the fix is non-trivial and management won't prioritize). Running on-call well is a culture problem. You need management to prioritize observa…
Or maybe page your managers, such that they can escalate the situation. They will be more aligned on solving the cultural problems if they get waked up too.
Show HN: I built an open-source tool to make on-call suck less
51–60 of 174 posts
Re: Show HN: I built an open-source tool to make on-call suck less
#52Shameless question tangential related to the topic. We are based in Europe and have the problem that some of us sometimes just forget we're on call or are afraid that we'll miss OpsGenie notifications. We're desparately looking for a hardware solution. I'd like something similar to the pagers of the past but at least here in Germany they don't really seem to exist anymore. Ideally I'd have a Bluetooth dongle that ale…
You can already enable silence/focus time bypass modes for apps like PagerDuty and such…
If you can’t develop some sense of responsibility to check if you’re on-call, frankly you have no business being in an on-call role.
No hardware will make you or your engineers more diligent. The only reason pagers made more sense than phones is/was because of protocol reasons NOT because it’s some separate device.
Re: Show HN: I built an open-source tool to make on-call suck less
#53In my current workplace (BigCo), we know exactly what's wrong with our alert system. We get alerts that we can't shut off, because they (legitimately) represent customer downtime, and whose root cause we either can't identify (lack of observability infrastructure) or can't fix (the fix is non-trivial and management won't prioritize). Running on-call well is a culture problem. You need management to prioritize observa…
I work on a team which runs hyper critical infra on all production machines at BigCo and have the same experience as you. The problem are not the alerts — the alerts actually are catching real problems — the problem is the following: 1. The team is understaffed so sometimes spending a few days root causing an alert is not prioritized 2. When alerts are root caused sometimes the work to fix the root cause is not prior…
Re: Show HN: I built an open-source tool to make on-call suck less
#54Re: Show HN: I built an open-source tool to make on-call suck less
#55Earlier quoted context omitted.
A candy bar cell phone, paid for by your employer and handed to whoever is on call. People who don't want it can just forward it to their phone.
in this case a satellite enabled candybar. the disaster recovery policy and budget should be applicable here. make sure its able to share xg and satellite tunnel for maximum value. ensure the reporting system is satellite enabled also. added points if its sending alerts 2 your handy byod. Disaster recovery is a big deal in 2024. All sorts of factors make satellite redundancy valuable in todays reality: Coworkers on a…
Re: Show HN: I built an open-source tool to make on-call suck less
#56In my current workplace (BigCo), we know exactly what's wrong with our alert system. We get alerts that we can't shut off, because they (legitimately) represent customer downtime, and whose root cause we either can't identify (lack of observability infrastructure) or can't fix (the fix is non-trivial and management won't prioritize). Running on-call well is a culture problem. You need management to prioritize observa…
> Running on-call well is a culture problem. You need management to prioritize observability (you can't fix what you can't show as being broken), then you need management to build a no-broken-windows culture (feature development stops if anything is broken). I was lucky enough to join a company where management does this. The managers were made to do this by experienced engineers who explained to them in no uncertain…
Sounds like managing up, i.e. doing IC workload and the manager's job. Hard pass.
Re: Show HN: I built an open-source tool to make on-call suck less
#57In my current workplace (BigCo), we know exactly what's wrong with our alert system. We get alerts that we can't shut off, because they (legitimately) represent customer downtime, and whose root cause we either can't identify (lack of observability infrastructure) or can't fix (the fix is non-trivial and management won't prioritize). Running on-call well is a culture problem. You need management to prioritize observa…
I work on a team which runs hyper critical infra on all production machines at BigCo and have the same experience as you. The problem are not the alerts — the alerts actually are catching real problems — the problem is the following: 1. The team is understaffed so sometimes spending a few days root causing an alert is not prioritized 2. When alerts are root caused sometimes the work to fix the root cause is not prior…
Re: Show HN: I built an open-source tool to make on-call suck less
#58Earlier quoted context omitted.
I work on a team which runs hyper critical infra on all production machines at BigCo and have the same experience as you. The problem are not the alerts — the alerts actually are catching real problems — the problem is the following: 1. The team is understaffed so sometimes spending a few days root causing an alert is not prioritized 2. When alerts are root caused sometimes the work to fix the root cause is not prior…
we as an industry need to have engineering management types realize that we cannot prioritize roadmap to the complete detriment of reliability
Re: Show HN: I built an open-source tool to make on-call suck less
#59Earlier quoted context omitted.
> Running on-call well is a culture problem. You need management to prioritize observability (you can't fix what you can't show as being broken), then you need management to build a no-broken-windows culture (feature development stops if anything is broken). I was lucky enough to join a company where management does this. The managers were made to do this by experienced engineers who explained to them in no uncertain…
> Communication with management is bidirectional, sometimes they need a lot of persuasion. Sounds like managing up, i.e. doing IC workload and the manager's job. Hard pass.
Re: Show HN: I built an open-source tool to make on-call suck less
#60Earlier quoted context omitted.
Or maybe page your managers, such that they can escalate the situation. They will be more aligned on solving the cultural problems if they get waked up too.
I have half-jokingly suggested that an out-of-hours page should cost the company $10k to incentivise actually fixing problems rather than releasing broken products. But I haven't thought of a way of getting around the perverse incentive to create bugs in order to get the $10k