I don't want to be relying on another flaky LLM for anything mission critical like this. Just fix the original problem, don't layer an LLM into it.
We wanted to provide that awareness because a lot of teams arent fully aware how bad the problem might be (on-calls change weekly, there might be a bunch of other issues)