- Loki
- Prometheus
- Grafana
- Incident.io
- GitHub for our codebase
- A custom MCP for our server inventory/nodes
Having all of them made it trivial for the LLM to start with an alert, search the codebase for the alert source and then review the Prom stack for additional info.
Automated a lot of the toil around digging through alerts and then making changes to fix issues.