Live data from Hacker News

Launch HN: Rootly (YC S21) – Manage Incidents in Slack

news.ycombinator.com

1–10 of 96 posts

Launch HN: Rootly (YC S21) – Manage Incidents in Slack

#1
Hi HN, Quentin and JJ here! We are co-founders at Rootly (https://rootly.com/), an incident management platform built on Slack. Rootly helps automate manual admin work during incidents like the creation of Slack channels, Jira tickets, Zoom rooms & more. We also help you get data on your incidents and help automate postmortem creation.

We met at Instacart, where I was the first SRE and JJ was on the product side owning ~20% GMV on the enterprise and last-mile delivery business. As Instacart grew from processing hundreds to millions of orders, we had to scale our infrastructure, teams, and processes to keep up with this growth. Unsurprisingly, this led to our fair share of incidents (e.g. checkout issues, site outages, etc.) and a lot of restless nights while on-call.

This was further compounded by COVID-19 and the first wave of lockdowns. We surged in traffic by 500% overnight as everyone turned to online grocery. This highlighted our need for a better incident management process as it stressed every element of it. Our manual ways of working in Slack, PagerDuty, Datadog, simply weren’t enough. At first, we figured this was an Instacart-specific problem but luckily realized it wasn’t.

A few things here. Our process lacked consistency. Depending on who was responding and their incident experience it varied greatly. Most companies after they declare an incident rely on a buried-away runbook like on Confluence/Google Docs to try and follow a lengthy checklist of steps. This is hard to find, difficult to follow accurately, slow, and stress inducing. Especially after you’ve been woken up to a page at 3 am. We started working on how to automate this.

Fast forward to today, companies like Canva, Grammarly, Bolt, Faire, Productboard, OpenSea, Shell use Rootly for their incident response. We think of ourselves as part of the post-alerting workflow. Tools like PagerDuty, Datadog act like a smoke alarm to alert you to an incident, which hand off to Rootly so we can orchestrate the actual response.

We’ve learned a lot along the way. We realized the majority of our customers use the same 6 (Slack, PagerDuty, Jira, Zoom, Confluence, Google Docs, etc.) tools, follow roughly the same incident response process (create incident → collaborate → write postmortem), but their process varies dramatically. The challenge in changing these processes is hard.

Our focus in the early days was build a hyper opinionated product to help them follow what we believe are the best practices. Now our product direction is focused on configuration and flexibility, how can we plug Rootly into your already existing way of working and automate it. This has helped our larger enterprise customers be successful with their current processes being automated.

Our biggest competition is not PagerDuty/Opsgenie (in fact 98% of our customers use them) or other startups. Its internal tooling companies have built out of necessity, often because tools like Rootly didn’t exist yet. Stripe (https://www.youtube.com/watch?v=fZ8rvMhLyI4) and GitLab (https://about.gitlab.com/handbook/engineering/infrastructure...) are good examples of this.

Our journey is just getting started as we learn more each day. Would love to hear any feedback on our product or anything you find frustrating about incident response today.

Leaving you with a quick demo: https://www.loom.com/share/313a8f81f0a046f284629afc3263ebff

Re: Launch HN: Rootly (YC S21) – Manage Incidents in Slack

#5
Congrats on the launch! This is really cool. I remember having to join several incidents and it was always a mess, especially people being left out, others being added who should not be there in the first place.

What happens when the incident is over? Where does all that data live and can there be some fancy data analytics that could potentially address bigger issues that keep reoccurring?

Re: Launch HN: Rootly (YC S21) – Manage Incidents in Slack

#6
post #4

Congrats on the launch. How is this different than FireHydrant?

Thank you!

Great question, There are quite a few differences, namely in our product design focus. We've taken a more configurable and flexible approach that focuses on plugging into a companies existing stack and their process. Often times we'll have customers send us their entire playbook on what they have now and ask us to automate that as a starting point (e.g. rename Slack channels to my Jira number for incidents, etc). We do this to hopefully reduce the amount of change required when a new tool is brought in. As a result we focus on features such as our Workflows engine that allows for this customization. Another big area of focus for us is unsurprisingly Slack, we think of the other areas of Rootly such as our Web platform to be the backend that powers this.

FH does a lot of things well and has great customers too. They have a sleek UI, strong security posture, and more. Their approach is more opinionated in guiding you through incident best practices. There is no wrong answer here as we hear plenty from customers that want both.

Re: Launch HN: Rootly (YC S21) – Manage Incidents in Slack

#7
I do t get the pricing on things like this and pingdom. This stuff seems like it should be cheap, like $5/user/mo. But everyone seems to go expensive.

There are other industries that are similar, but this always stood out to me as an industry where the pricing never felt right to me.

Re: Launch HN: Rootly (YC S21) – Manage Incidents in Slack

#8
post #6
post #4

Congrats on the launch. How is this different than FireHydrant?

Thank you! Great question, There are quite a few differences, namely in our product design focus. We've taken a more configurable and flexible approach that focuses on plugging into a companies existing stack and their process. Often times we'll have customers send us their entire playbook on what they have now and ask us to automate that as a starting point (e.g. rename Slack channels to my Jira number for incidents…

Thank you.

Re: Launch HN: Rootly (YC S21) – Manage Incidents in Slack

#9
post #5

Congrats on the launch! This is really cool. I remember having to join several incidents and it was always a mess, especially people being left out, others being added who should not be there in the first place. What happens when the incident is over? Where does all that data live and can there be some fancy data analytics that could potentially address bigger issues that keep reoccurring?

Many thanks!

So glad you asked. Once the incident is in a resolved, we'll prompt you to edit your postmortem. This can be done inside of Rootly but most commonly we'll auto-generate a Confluence or Google Doc. Here you'll have all your incident metadata, template to fill out, but most importantly your incident timeline (no copy-paste required).

From there we can help you do things like automatically scheduling your postmortem meeting with everyone that was involved.

We also want to help you improve your process and response. We'll prompt anyone involved in the incident for feedback (they can submit anonymously too) and collect important metrics.

There are top line metrics like incident count, MTTX, outstanding action items but also finer grained ones which is what I think you might be hinting at. For example you can visualize automatically what services are being impacted the most. That might be an indicator as an area to focus on more.

We try to keep the garden walls on the product quite low, allowing you to export any of the data out of Rootly and into your own analytics engines.

Post reply on HN