I love the AWS postmortems for their simplicity, timestamps, and insight into internal AWS systems: https://aws.amazon.com/premiumsupport/technology/pes/
Famous outages along with deep postmortems?
11–13 of 13 posts
Re: Famous outages along with deep postmortems?
#12Microsoft was about to spend $500 million on a blitz ad campaign called Five Nines, for 99.999% uptime re NT 5, was it? 2002ish.
They crashed the microsoft.com cluster only days before, sending certain accepted metrics re: uptime from 99.999% to 97.312%.
The cluster crash was caused by some errant out-of-band JavaScript being published to a live MSCOM cluster. The postmortem was not in-depth, it was a cremation. Burn and hide the body.
I thought it odd that all those involved ended up at AWS shortly thereafter, including the executive whose head rolled right out of 1 Microsoft Way.
Those involved owe Dave Cutler an apology, with or without the conspiracy intact.
Re: Famous outages along with deep postmortems?
#13Perhaps not famous, but Bryan Cantrill, who gives my favorite talks, has an interesting and funny talk on one of the Joyent outages: https://youtu.be/30jNsCVLpAE