Post-incident review on the Atlassian April 2022 outage
1–10 of 157 posts
Re: Post-incident review on the Atlassian April 2022 outage
#2I’ve been there with much less significant incidents when a “routine” change turned into a potentially resume generating event. It’s not fun.
Ultimately the responsibility is with the organization that made it possible for an event of that scale to happen rather than the individual person who happened to trigger it but that doesn’t make it feel any better.
Re: Post-incident review on the Atlassian April 2022 outage
#3I can only imagine the way the person who pulled the trigger on the deletion script felt the moment they realized what had happened. I’ve been there with much less significant incidents when a “routine” change turned into a potentially resume generating event. It’s not fun. Ultimately the responsibility is with the organization that made it possible for an event of that scale to happen rather than the individual pers…
Re: Post-incident review on the Atlassian April 2022 outage
#4Yup that deletes something… anyway…
> Establish universal "soft deletes" across all systems.
It’s just easier that way to observe what might happen.
Re: Post-incident review on the Atlassian April 2022 outage
#51. Establish universal "soft deletes" across all systems.
2. Better DR for multi-site, multi-product incidents.
3. Fix their incident-management process for large-scale incidents.
4. Fix their incident communications.
Regarding #4 in particular:
"Rather than wait until we had a full picture, we should have been transparent about what we did know and what we didn't know. Providing general restoration estimates (even if directional) and being clear about when we expected to have a more complete picture would have allowed our customers to better plan around the incident....
[In the future], we will acknowledge incidents early, through multiple channels. We will release public communications on incidents within hours. To better reach impacted customers, we will improve the backup of key contacts and retrofit support tooling to enable customers... [to make emergency] contact with our technical support team."
Re: Post-incident review on the Atlassian April 2022 outage
#6So how would one have a clean "partial backup" strategy if something similar would happen in his company?
Re: Post-incident review on the Atlassian April 2022 outage
#7I can only imagine the way the person who pulled the trigger on the deletion script felt the moment they realized what had happened. I’ve been there with much less significant incidents when a “routine” change turned into a potentially resume generating event. It’s not fun. Ultimately the responsibility is with the organization that made it possible for an event of that scale to happen rather than the individual pers…
As someone who observed this particular incident from the inside (holy shit-balls have the last 3 weeks not been fun), one of the few positive elements of it has been the universal and effectively instinctual agreement internally that it was a massive screw-up in the system that we all have to own, rather than one or several individual screw-ups that need to be put at the feet of individuals.
Re: Post-incident review on the Atlassian April 2022 outage
#8Would be super interested in a technical writeup on how they do this.
Re: Post-incident review on the Atlassian April 2022 outage
#9Re: Post-incident review on the Atlassian April 2022 outage
#10Soft deletion always feels at odds with privacy-related "right to have data deleted" laws. Would be super interested in a technical writeup on how they do this.