Live data from Hacker News

Leveraging AI for efficient incident response

engineering.fb.com

41–50 of 58 posts

Re: Leveraging AI for efficient incident response

#41

AI 1: This user is suspicious, lock account User: Ahh, got locked out, contact support and wait AI 2: The user is not suspicious, unlock account User: Great, thank you AI 1: This account is suspicious, lock account

Luckily I subscribe to my own consumer AI service to automate all this for me. To paraphrase The Simpsons: "AI: the cause of and solution to all life's problems."

Re: Leveraging AI for efficient incident response

#42
post #11

We've shifted our oncall incident response over to mostly AI at this point. And it works quite well. One of the main reasons why this works well is because we feed the models our incident playbooks and response knowledge bases. These playbooks are very carefully written and maintained by people. The current generation of models are pretty much post-human in following them, performing reasoning and suggesting mitigati…

Curious if you explored any external tools before building in-house? Looking to do something similar at my company

Re: Leveraging AI for efficient incident response

#44
post #11

We've shifted our oncall incident response over to mostly AI at this point. And it works quite well. One of the main reasons why this works well is because we feed the models our incident playbooks and response knowledge bases. These playbooks are very carefully written and maintained by people. The current generation of models are pretty much post-human in following them, performing reasoning and suggesting mitigati…

That's great to hear. What is your current tool chain in the effort? Do you have a structure for Playbooks and KBs you would recommend

Re: Leveraging AI for efficient incident response

#46
post #22
post #11

We've shifted our oncall incident response over to mostly AI at this point. And it works quite well. One of the main reasons why this works well is because we feed the models our incident playbooks and response knowledge bases. These playbooks are very carefully written and maintained by people. The current generation of models are pretty much post-human in following them, performing reasoning and suggesting mitigati…

I'm really curious to hear more about what kind of thing is covered in your playbooks. I've often heard and read about the value of playbooks, but I've yet to see it bear fruit in practice. My main work these past few years has been in platform engineering, and so I've also been involved in quite a few incidents over that time, and the only standardized action I can think of that has been relevant over that time is c…

Playbooks that I've found value in: - Generic application version SLI comparison. The automated version of this is automated rollbacks (Harness supports this out of the box, but you can certainly find other competitors or build your own) - Database performance debugging - Disaster recovery (bad db delete/update, hardware failure, region failure)

In general, playbooks are useful for either common occurences that happen frequently (ie every week we need to run a script to fix something in the app) or things that happen rarely but when they do happen need a plan (ie disaster recovery)

Re: Leveraging AI for efficient incident response

#48
post #11

We've shifted our oncall incident response over to mostly AI at this point. And it works quite well. One of the main reasons why this works well is because we feed the models our incident playbooks and response knowledge bases. These playbooks are very carefully written and maintained by people. The current generation of models are pretty much post-human in following them, performing reasoning and suggesting mitigati…

Meanwhile, we’ve tried AI products just for assigning incidents and are forced to turn them off because of how shitty of a job they do.

Re: Leveraging AI for efficient incident response

#49
post #11

We've shifted our oncall incident response over to mostly AI at this point. And it works quite well. One of the main reasons why this works well is because we feed the models our incident playbooks and response knowledge bases. These playbooks are very carefully written and maintained by people. The current generation of models are pretty much post-human in following them, performing reasoning and suggesting mitigati…

I have found it rare that an organization has incident "playbooks that are very carefully written and maintained"

If you already have those, how much can an AI add? Or conversely, not surprising that it does well when it's given a pre-digested feed of all the answers in advance.

Re: Leveraging AI for efficient incident response

#50
post #35
post #23

Earlier quoted context omitted.

Sounds like you're projecting your own laziness and shortcomings on others. This is a tool that seems really helpful considering the alternative is 0%.

Personal insults aside, "seems" requires no evaluation if the success rate is outside what could be considered a sane confidence interval on trust. I would literally be fired if I implemented this tool.

Calling things 'shit' and 'crap,' and then claiming that the authors actually feel the same but can't say it, is ridiculous and undermines any authority you think you have.
Post reply on HN