Live data from Hacker News

Launch HN: Rootly (YC S21) – Manage Incidents in Slack

news.ycombinator.com

81–90 of 96 posts

Re: Launch HN: Rootly (YC S21) – Manage Incidents in Slack

#81
post #24

How is Blameless doing?

We just picked them up across an engineering/tech org of ~350 to do precisely what they describe here.

PagerDuty for notifications and on-call rotations, Datadog for monitoring, Slack for communication in-the-moment, Google Docs for post-mortem documentation; Blameless as the glue and automation that takes away a lot of the incidental mental overhead of communicating and documenting while the incident is happening.

Super encouraging to see competition, though. A former teammate turned me on to https://how.complexsystems.fail/ and I'm willing to believe that in a complex enough system, the closest we will get to understanding how it actually works is during/after incident response.

Re: Launch HN: Rootly (YC S21) – Manage Incidents in Slack

#82
post #31

I’ve always wondered about building a startup on another startup’s back. What happens if they cut you off? Is getting bought up by Slack the end goal here? Seems like a big risk, one whim at Slack and you’re toast.

We have seen Slack start investing in this area with their own Workflow Builder they announced last year. One of the big use cases they highlighted was incident response. We haven't ran into any customers trying to leverage that just yet though as still a lot of heavy lifting required. IMO what makes Slack so power is their app ecosystem. We aren't too worried about them shutting that down or competing with us. We se…

I'm not familiar with Workflow Builder. Does it have a separate data retention scheme? Slack seems to have a big problem in that information quickly ages out, the search is pretty bad, and, in some cases, the data is actually ephemeral and will just disappear. Incident response is one of the categories that I want to preserve. Does Rootly address any of these problems?

OK, I see at the end of the demo that there is a chat transcript, so that's useful. Does it differentiate between incidents if there are multiple active incidents? Where is that archived stored?

Re: Launch HN: Rootly (YC S21) – Manage Incidents in Slack

#84
post #78

I will echo the other comments on no upfront pricing. Even though this could be potentially useful for my team, I won't "contact" you for pricing. I am sick of having to deal with salespeople who want to know a ton of info about your business so they can gouge every last cent out of you and then some. I would gladly pay a little extra just to have clear pricing and sign up with a credit card. I have got an engineerin…

I wish it was also standard to provider a small sandbox environment where I can go setup things and play around. I don't want to book a demo where I'll again be hooked up with someone from sales for a long presentation. For someone with a few million in funding, there should really be no excuse to not invest in this.

Totally agree, we do offer a 14-day free trial here if you want to give it a go: https://rootly.com/users/sign_up.

This feedback is helpful, I think we can make that a bit more obvious on our website!

Re: Launch HN: Rootly (YC S21) – Manage Incidents in Slack

#85
post #49

Earlier quoted context omitted.

Thank you for the feedback, a calculator is a good suggestion! We realize there are a fair number of people that will be turned away by this, we'll see what we can do for a better middle ground.

I’d also suggest you have a free plan for very small teams. You can already see how many slack members they have. Make your tool so people just make it their default and then as the team grows the naturally start paying you up the tiers.

Thank you for the feedback.

Actually we did offer a free plan for up to 5 users and even 10 at one point. What we've found is the collaboration overhead during incidents at companies at that scale wasn't too useful and pivoted away. Instead we offer a 14-day trial so customers can get their feet wet without contacting us: https://rootly.com/users/sign_up.

Re: Launch HN: Rootly (YC S21) – Manage Incidents in Slack

#86
post #24

How is Blameless doing?

We just picked them up across an engineering/tech org of ~350 to do precisely what they describe here. PagerDuty for notifications and on-call rotations, Datadog for monitoring, Slack for communication in-the-moment, Google Docs for post-mortem documentation; Blameless as the glue and automation that takes away a lot of the incidental mental overhead of communicating and documenting while the incident is happening. S…

One of my favourite sites!

And that is great, we know plenty of happy Blameless customers, they're certainly one of the better ones in the market we compete against.

Re: Launch HN: Rootly (YC S21) – Manage Incidents in Slack

#87
post #82
post #31

Earlier quoted context omitted.

We have seen Slack start investing in this area with their own Workflow Builder they announced last year. One of the big use cases they highlighted was incident response. We haven't ran into any customers trying to leverage that just yet though as still a lot of heavy lifting required. IMO what makes Slack so power is their app ecosystem. We aren't too worried about them shutting that down or competing with us. We se…

I'm not familiar with Workflow Builder. Does it have a separate data retention scheme? Slack seems to have a big problem in that information quickly ages out, the search is pretty bad, and, in some cases, the data is actually ephemeral and will just disappear. Incident response is one of the categories that I want to preserve. Does Rootly address any of these problems? OK, I see at the end of the demo that there is a…

Rootly would address that problem, we keep a database of all of your incidents and metadata (impact, timeline, participants, metrics, etc.) on our Web platform separate from Slack. You can customize a data retention policy with us if you want but it's helpful to be able to quickly search for similar incidents without trying to find it in Slack channels.

It does differentiate between incidents if there are multiple too. We'll even warn you if you're opening an incident and another one that could be related is also active to avoid duplications.

And of course we keep the garden walls low on the product. You can export any of this data out via CSV, JSON, API or via our integrations (Airtable, Google Sheets, Looker, etc.).

Re: Launch HN: Rootly (YC S21) – Manage Incidents in Slack

#88

I'm very curious, why base everything on Slack? There's no denying it's super popular (and my preferred choice for collab) but the majority of businesses don't use it, and it indirectly adds cost to your product; if I want to use (and pay for) Rootly I also have to pay for Slack (which is expensive). And to echo everyone else here; not having up front pricing is a big red flag. You've done well to attract the large c…

The messaging for "managing incidents on Slack" is a lot easier to understand and digest we found than something like "all-in-one incident management platform".

We have an entire Web platform that performs the same functions as Slack but also can configure custom Workflows, integrations, view metrics, service catalog, and more.

And thank you for the feedback on pricing, we have a bit of work to do here to make it easier to understand the financial cost before investing time to chat with us!

Re: Launch HN: Rootly (YC S21) – Manage Incidents in Slack

#89
post #33
post #24

How is Blameless doing?

From our experience we still see them in a number of deals. I think their focus has shifted more towards SLOs and Microsoft Teams though, areas where we aren't investing in right now.

Cool, thanks for this view.

I'm also intrigued by the text in this launch announcement:

> Our focus in the early days was build a hyper opinionated product to help them follow what we believe are the best practices. Now our product direction is focused on configuration and flexibility, how can we plug Rootly into your already existing way of working and automate it. This has helped our larger enterprise customers be successful with their current processes being automated.

As I have gotten more experience managing complex incidents I've come around to the idea that having a standard process you follow for big issues is somewhat more important than what the process really is.

I loved the PagerDuty response documentation ( https://response.pagerduty.com/ ) not so much because of the specifics but because it suggests they have a culture where there is a well-understood protocol they always try to follow for big problems.

I think about archery and "shot grouping" - once you learn to always land in the same place, you can move your aim to start landing somewhere else.

A number of the things that I see as valuable incident management involve having responders with a shared set of priorities. Tooling can influence how easy/hard some of these things are but it's really up to the people to do things like:

* Actually finding and fixing the problem and being sure the fix worked

* Clearly communicating the current user impact to the people who care

* Figuring out who the right responders are, and getting them in the room quickly

* Making one production change at a time with the incident coordinator's signoff, so you know which one helped and when it happened

* Helping the rest of the organization learn from what happened (you may not know what there is to learn)

Do you see room for the tooling company to also provide best-practices training, mentorship, or other kinds of support? That stuff scales less well than a web app but is arguably more important to changing a company's culture in a way that gets better user outcomes.

Post reply on HN