Earlier quoted context omitted.
You might be interested in https://response.pagerduty.com/ , PagerDuty's major incident response process documentation - a good starting point for that red binder. Having been in the ringmasters seat for major incidents ranging from "relatively routine" to "it's all on fire", and had a ringside seat for a cloud provider outage of comparable magnitude to this one - it still fascinates me how creative solutions can get…
Anyone know of other public resources like the one from PagerDuty? The SRE book and workbook at https://landing.google.com/sre/books/ have some details, but curious if there are others people would recommend.
I particularly like Gene Kranz "Failure is not an Option". It is more background but it works. In general, it is not crazy hard. You get the roles, you distribute them. Someone can have multiple roles that depends on the size of the incident.
The usual roles i differentiate are Point (think of it as IC if you want), Comms and Logs