Live data from Hacker News

Post-incident review on the Atlassian April 2022 outage

atlassian.com

1–10 of 157 posts

Re: Post-incident review on the Atlassian April 2022 outage

#2
I can only imagine the way the person who pulled the trigger on the deletion script felt the moment they realized what had happened.

I’ve been there with much less significant incidents when a “routine” change turned into a potentially resume generating event. It’s not fun.

Ultimately the responsibility is with the organization that made it possible for an event of that scale to happen rather than the individual person who happened to trigger it but that doesn’t make it feel any better.

Re: Post-incident review on the Atlassian April 2022 outage

#3

I can only imagine the way the person who pulled the trigger on the deletion script felt the moment they realized what had happened. I’ve been there with much less significant incidents when a “routine” change turned into a potentially resume generating event. It’s not fun. Ultimately the responsibility is with the organization that made it possible for an event of that scale to happen rather than the individual pers…

As someone who observed this particular incident from the inside (holy shit-balls have the last 3 weeks not been fun), one of the few positive elements of it has been the universal and effectively instinctual agreement internally that it was a massive screw-up in the system that we all have to own, rather than one or several individual screw-ups that need to be put at the feet of individuals.

Re: Post-incident review on the Atlassian April 2022 outage

#4
>The script that was executed followed our standard peer-review process that focused on which endpoint was being called and how. It did not cross-check the provided cloud site IDs to validate whether they referred to the Insight App or to the entire site, and the problem was that the script contained the ID for a customer's entire site.

Yup that deletes something… anyway…

> Establish universal "soft deletes" across all systems.

It’s just easier that way to observe what might happen.

Re: Post-incident review on the Atlassian April 2022 outage

#5
I can't say that I've ever been a fan of Atlassian or their products, but this blog post makes it sounds like they've at least learned the right lessons from this:

1. Establish universal "soft deletes" across all systems.

2. Better DR for multi-site, multi-product incidents.

3. Fix their incident-management process for large-scale incidents.

4. Fix their incident communications.

Regarding #4 in particular:

"Rather than wait until we had a full picture, we should have been transparent about what we did know and what we didn't know. Providing general restoration estimates (even if directional) and being clear about when we expected to have a more complete picture would have allowed our customers to better plan around the incident....

[In the future], we will acknowledge incidents early, through multiple channels. We will release public communications on incidents within hours. To better reach impacted customers, we will improve the backup of key contacts and retrofit support tooling to enable customers... [to make emergency] contact with our technical support team."

Re: Post-incident review on the Atlassian April 2022 outage

#6
If I understand correctly, since they deleted only a small subset of all their customers, they could not restore a clean backup of those customers without losing data from other customers not impacted by the outage.

So how would one have a clean "partial backup" strategy if something similar would happen in his company?

Re: Post-incident review on the Atlassian April 2022 outage

#7

I can only imagine the way the person who pulled the trigger on the deletion script felt the moment they realized what had happened. I’ve been there with much less significant incidents when a “routine” change turned into a potentially resume generating event. It’s not fun. Ultimately the responsibility is with the organization that made it possible for an event of that scale to happen rather than the individual pers…

As someone who observed this particular incident from the inside (holy shit-balls have the last 3 weeks not been fun), one of the few positive elements of it has been the universal and effectively instinctual agreement internally that it was a massive screw-up in the system that we all have to own, rather than one or several individual screw-ups that need to be put at the feet of individuals.

Well, the one constant of IT is 'shit happens'; mark this one down as something interesting you've seen.

Re: Post-incident review on the Atlassian April 2022 outage

#10
post #8

Soft deletion always feels at odds with privacy-related "right to have data deleted" laws. Would be super interested in a technical writeup on how they do this.

It shouldn’t be. These laws at least have the nuance to understand that data can’t be immediately deleted from Backups and that in such instances where deletes are complicated the customer is notified.
Post reply on HN