Live data from Hacker News

Show HN: I built an open-source tool to make on-call suck less

github.com

41–50 of 174 posts

Re: Show HN: I built an open-source tool to make on-call suck less

#41
post #12
post #11

In my current workplace (BigCo), we know exactly what's wrong with our alert system. We get alerts that we can't shut off, because they (legitimately) represent customer downtime, and whose root cause we either can't identify (lack of observability infrastructure) or can't fix (the fix is non-trivial and management won't prioritize). Running on-call well is a culture problem. You need management to prioritize observa…

I completely agree that technical tools cannot fix culture problems. However, one of the things that I noticed in my previous companies was that my management chain wasn't even aware that the problem was this bad. We also wanted to add better reporting (like the alert analytics) so that people have more visibility into the state of alerts + on-call load on engineers. What strategies have worked well for you when it c…

Show them the costs! Wasted time, wasted resources, wasted money. Show the waste and come with the plan to reduce the waste. Alerts, on-calls and tests are all waste reduction.

"We're paying down our technical debt"

Re: Show HN: I built an open-source tool to make on-call suck less

#42
post #40

Earlier quoted context omitted.

That’s true. But technical tools can help you highlight culture problems so that they’re easier to to discuss and fix. It’s been a minute since I’ve had to process exactly the kind of on-call/alert problem we’re discussing here, but this does feel like the kind of tool that would help sell the kinds of management/culture changes necessary to really improve things, if not fix all of them.

Switching tools, or adopting new (unproven) ones doesn't address or fix the communication issue. The existing tools mentioned can show the metrics. Management needs an education - and that is part of the engineering job.

> Management needs an education - and that is part of the engineering job

Isn’t that bizarre? In all my years as an engineer I can count the number of managers that went to learn about engineering by themselves, on one hand.

It’s literally their job, but somehow they feel they can do it without understanding it.

Re: Show HN: I built an open-source tool to make on-call suck less

#43
post #3

every time I see notifications in Slack / Telegram it makes me depressed. Text messengers were not designed for this. If you get the "something is wrong" alert it becomes part of history, it won't re-alert you if it's still present. And if you have more than one type of alert it will be lost in history I guess alerts to messengers are OK as long it's only a couple manually created ones, and there should be a graphica…

Why? We send alerts to Slack and Pagerduty. Slack is to help everyone who might be working, PagerDuty alerts the persons who are actually in charge of working on it.

Yeah, I think it’s convenient. We use email, but for the same thing. If I inadvertedly break something, I’ll have an email in my inbox 5 minutes later.

Re: Show HN: I built an open-source tool to make on-call suck less

#44

Earlier quoted context omitted.

Try Netherlands. We're Microsoft land over here. Pretty much everyone is on Azure and Teams. It's mostly startups and hip small companies that use Slack.

Startups, hip small companies, tech product based companies. Most non tech product based or enterprise banks in NL are on Teams

I really feel like the world would be a better place if it was illegal to bundle Teams like this…

Re: Show HN: I built an open-source tool to make on-call suck less

#46
I feel like this would be a great tool for people who have had a much better experience of On Call than I have had.

I once worked for a string of businesses that would just send everything to on call unless engineers threatened to quit. Promised automated late night customer sign ups? Haven't actually invested in the website so that it can do that? Just make the on call engineer do it. Too lazy to hire off shore L1 technical support? Just send residential internet support calls to the the On Call engineer! Sell a service that doesn't work in the rain? Just send the on call guy to site every time it rains so he can reconfirm yes, the service sucks. Basic usability questions that could have been resolved during business hours? Does your contract say 24/7 support? Damn, guess thats going to On Call.

Shit even in contracting gigs where I have agreed to be "On Call" for severity 1 emergencies, small business owners will send you things like service turn ups or slow speed issues.

Re: Show HN: I built an open-source tool to make on-call suck less

#47

What you've come up with looks helpful (and may have other applications as someone else noted), but you know what also makes on-call suck less? Getting paid for it, in $ and/or generous comp time. :-) https://betterstack.com/community/guides/incident-management... Also helpful is having management that is responsive to bad on-call situations and recognizes when capable, full-time around-the-clock staffing is really n…

I guess 7-Eleven management trainees know that their company is just as replacable for their employees as their employees are to them.

Re: Show HN: I built an open-source tool to make on-call suck less

#48
post #3

every time I see notifications in Slack / Telegram it makes me depressed. Text messengers were not designed for this. If you get the "something is wrong" alert it becomes part of history, it won't re-alert you if it's still present. And if you have more than one type of alert it will be lost in history I guess alerts to messengers are OK as long it's only a couple manually created ones, and there should be a graphica…

I don't think Slack (or similar) should be a primary alert mechanism.

But, if alarms are configured in a clean way, ideally your team is getting some warnings and such there and then if there's an alert that needs to actually page, it sends that to PagerDuty or whatever platform you use along with another message to Slack.

Re: Show HN: I built an open-source tool to make on-call suck less

#49
post #42
post #40

Earlier quoted context omitted.

Switching tools, or adopting new (unproven) ones doesn't address or fix the communication issue. The existing tools mentioned can show the metrics. Management needs an education - and that is part of the engineering job.

> Management needs an education - and that is part of the engineering job Isn’t that bizarre? In all my years as an engineer I can count the number of managers that went to learn about engineering by themselves, on one hand. It’s literally their job, but somehow they feel they can do it without understanding it.

I don't think it is bizarre. I see lots of MBAs running things. They don't have the engineering background, they have the "resources management" background.

I think engineer brings the numbers to management to decide course.

I prefer the situation where the CTO has no MBA and worked their way up - but that is uncommon IME.

So, in many orgs, engineer puts their comms hat on an presents a solid case.

The engineer who can communicate well, and show the metrics is typically the one who can get promoted to the decision maker role. First from the bottom up, then as a great leader

Re: Show HN: I built an open-source tool to make on-call suck less

#50
post #3

every time I see notifications in Slack / Telegram it makes me depressed. Text messengers were not designed for this. If you get the "something is wrong" alert it becomes part of history, it won't re-alert you if it's still present. And if you have more than one type of alert it will be lost in history I guess alerts to messengers are OK as long it's only a couple manually created ones, and there should be a graphica…

I would expect anything notifying via Slack or text would have an accompanying incident ticket in the system of record.

We had a rule in my team (before a management change that blew it all to shit) that we don’t use email or messaging for monitoring. Everything goes into the SOR. Once it’s in the SOR, if people want emails, texts, or whatever, it let them know there is work to do, that’s up to the team. Others would make dashboards… lots of options once it’s in the system, and nothing gets lost.

For example, I went from a team that looked at tickets all day to one that mainly worked on user stories in Jira. Because no one was looking at the incidents in the SOR, things were getting missed. I wrote something to check for incident tickets assigned to our team every hour, and it would post them in our team chat so people knew there was work to do. Then once per day, it would post everything still unassigned, so if something was lost on that hourly post, it would annoy everyone once per day until it was assigned/resolved. It worked out decently well. If there was a lot of stuff, it would post a message to have someone actually login to the SOR and look at all our tickets. I would sometimes use the standup to assign stuff out and get some attention on it, if things were getting bad.

Post reply on HN