Live data from Hacker News

Post-incident review on the Atlassian April 2022 outage

atlassian.com

71–80 of 157 posts

Re: Post-incident review on the Atlassian April 2022 outage

#71

Earlier quoted context omitted.

I tend to agree. However, there's one point that makes me skeptical: there are no organizational changes, or changes to leadership, or anything in that direction. This sounds like "the tech guys screwed up, culture and management is fine here". Which it might be, or it might not. I would have loved to see 5. We will stop pushing customers so hard towards using our cloud for one, but that wouldn't be convenient for At…

Note that the ToS also forbid Cloud users from “disseminating information” about “the performance of the products”. So you can’t say it’s slow or unperforming.

Link for reference: https://www.atlassian.com/legal/cloud-terms-of-service

> Except as otherwise expressly permitted in these Terms, you will not:

> ... (i) publicly disseminate information regarding the performance of the Cloud Products;

Re: Post-incident review on the Atlassian April 2022 outage

#72

Earlier quoted context omitted.

Note that the ToS also forbid Cloud users from “disseminating information” about “the performance of the products”. So you can’t say it’s slow or unperforming.

Well that's some Oracle-tier shit. (ref: https://danluu.com/anon-benchmark/ )

Not just benchmarks, Oracle also sues (or threatens) security researchers, vendors, and of course customers for less.

Re: Post-incident review on the Atlassian April 2022 outage

#73

Implementing soft deletes is a lessons every developer learns early in their career. The fact that Atlassian did not implement that in their cloud, is mind boggling. Great case study of a monumental fuck up!

Soft deletes are great, but are probably not sufficient to meet the GDPR's "right to be forgotten"

Re: Post-incident review on the Atlassian April 2022 outage

#74
post #73

Implementing soft deletes is a lessons every developer learns early in their career. The fact that Atlassian did not implement that in their cloud, is mind boggling. Great case study of a monumental fuck up!

Soft deletes are great, but are probably not sufficient to meet the GDPR's "right to be forgotten"

There's no reason you can't have both. Use soft deletes for everything except for a formal GDPR right to be forgotten request (or any other compliance situation).

Re: Post-incident review on the Atlassian April 2022 outage

#75
While the post-mortem is thorough, it misses key details on what companies experienced who were unlucky enough to be caught out by this outage. For example, it fails to mention how impacted customers lost access to certain Atlassian services for up to ~2 weeks: JIRA, Confluence, OpsGenie. But not others like Trello or BitBucket.

Of these, losing access to OpsGenie for this long was a massive problem, dwarfing most other systems. OpsGenie is like PagerDuty in the Atlassian world.

I spoke to several engineers at impacted companies who could not believe their incident management system was “deleted” and had no ETA on when it would be back, or Atlassian could not prioritise restoring this critical system ASAP. JIRA and Confluence being down was trouble enough, but those systems being down for some time was things most teams worked around. However, suddenly flying blind, with no pager alerting for their own systems? That is not acceptable for any decent company.

Most I talked with moved rapidly to an alternative service, building up oncall rosters from memory and emails - as Confluence which stored these details was also down. Imagine being a billion dollar company suddenly without pager system: and no ETA on when that system would be back, your vendor not responding to your queries.

I talked to engineers at such a company and it was a long night to move rapidly over to PagerDuty. It would be another 7 days they could get through to a human at Atlassian. By that time, they were a lost customer for this product. Ironically, this company moved to OpsGenie a few years before from PagerDuty because OpsGenie was cheaper and they were on so many Atlassian services already.

The post-mortem has no actions on prioritising services like OpsGenie in reliability or restoration, which is a miss. I can’t tell if Atlassian staff are unaware of the critical nature of this system or if they treat all their products - including paging systems - as equals in terms of SLAs on principle.

Worth keeping in mind when choosing paging vendors - some might recognise these systems are more critical ones than others.

I wrote about this outage from the viewpoint of the customers as it entered its 10th day and it was discussed in HN, with comments from people impacted by the outage. [1]

[1] https://news.ycombinator.com/item?id=31015813

Re: Post-incident review on the Atlassian April 2022 outage

#76
post #70

I sincerely hope all of the people who worked on the recovery effort are okay, and are being well supported and strongly encouraged to tend their mental health. I have no personal investment in Atlassian products—if anything, unrelated to this or any incident, I could happily never use them again. But the people who work there are human, and I know what kind of a toll a protracted recovery effort can take. I know it…

As someone who's worked on similar situations, I don't expect any pat on the back, but see it as my responsibility to make sure it doesn't happen in the first place, and when it does, my responsibility to fix it without complaint. Some Atlassian customers might have had much more severe (mental health and other) problems than the Atlassian staff, so I think it can be perceived as mildly solipsistic to be praising the…

This is a fascinating response. I’ll address the last part first, because I don’t want it to get lost.

> I know you didn't say anything about forgetting about the customers, but affected customers may perceive it that way

No, I didn’t mean to suggest or imply this, and hope it won’t be taken this way by anyone. It was perhaps a mistake taking it as read that obviously their customers were also harmed. But I do think it’s a reasonable point to make more explicit, and I do agree that it likely had a similar impact on customers to the experience I described. I’ll repeat that no one should experience that.

That said…

> As someone who's worked on similar situations, I don't expect any pat on the back, but see it as my responsibility to make sure it doesn't happen in the first place, and when it does, my responsibility to fix it without complaint.

This seems like you’re addressing something entirely outside the actual content of my comment, and perhaps projecting your own priorities onto it.

First of all, I was in no way suggesting anyone get special reward. And I was in no way referencing any Atlassian employee’s complaints nor airing my own. I mean it sincerely that I hope they are okay and that their mental health needs are being respected. It was entirely a statement of human compassion, and an observation that it’s one which can go unstated/understated in these discussions.

Secondly, I added my own experience as a personal reflection on the toll it can take. It’s not easy to say in a public forum that I’ve suffered years of chronic pain after addressing an incident. I added this for context because I think it is easy for people to dismiss the impact serious incidents have on the people responsible to them.

A side note before I get to thirdly: I think this is also true for people in many careers where incident response is a primary job responsibility. Sometimes it prompts explicit acknowledgement, often it does not. I just think the world would be better off if more people’s legitimate pain and challenges were acknowledged.

Third point: I took special care above to say “responsible to”, not “responsible for” or simply “responsible”. I’m certain, given Atlassian’s size and the impact of this outage, that many of the people involved in recovery efforts played no role in the mistakes leading to the outage.

In my own anecdote, I played no role in causing the incident. I have to walk a fine professional line here because I have no intention or desire to criticize anyone else involved. I think I can reasonably say this. I did what you described:

> see it as my responsibility to make sure it doesn't happen in the first place, and when it does, my responsibility to fix it without complaint.

Even so, I’m experiencing chronic pain years later. And now I will complain: it sucks being in pain every day for years. I wouldn’t go back and do anything differently as an IC, except perhaps to tell past me when to slow down and that I have more ability to affect short term prioritization than I once realized.

Lastly,

> praising the staff

I sincerely hope this hasn’t had the negative impact on you which you’re describing as hypothetical for Atlassian customers generally. If it has, I hope you’ll hear me when I say any sharp tone in this response is not from lack of solidarity.

But if it hasn’t and you’re replying only out of work ethic and blame-placing: kindly go re-read my comment and recognize that the only praise I expressed was for the humanity of human people whose experience I wish to include in the conversation.

Re: Post-incident review on the Atlassian April 2022 outage

#77

Implementing soft deletes is a lessons every developer learns early in their career. The fact that Atlassian did not implement that in their cloud, is mind boggling. Great case study of a monumental fuck up!

So many opportunities missed to avoid this! Look at one ID and ensure it is what you expect it to be. Run the script in a dryrun mode. Run the script for 1 customer. Probably more!

This was addressed in the write up (it’s very long, so missing it is easy). They ran the script against 30 accounts first to verify it worked, and it did, because the list of 30 ids they tested against came from a different source than the other 750ish. It’s a shitty mistake to make but I’m certain I’ve made similar ones.

Re: Post-incident review on the Atlassian April 2022 outage

#78
post #73

Earlier quoted context omitted.

Soft deletes are great, but are probably not sufficient to meet the GDPR's "right to be forgotten"

There's no reason you can't have both. Use soft deletes for everything except for a formal GDPR right to be forgotten request (or any other compliance situation).

Soft deletes are just the first step. After some period of time you can automatically purge or manually purge. Which is what they’ve supposedly committed to doing.

Re: Post-incident review on the Atlassian April 2022 outage

#79
post #70

I sincerely hope all of the people who worked on the recovery effort are okay, and are being well supported and strongly encouraged to tend their mental health. I have no personal investment in Atlassian products—if anything, unrelated to this or any incident, I could happily never use them again. But the people who work there are human, and I know what kind of a toll a protracted recovery effort can take. I know it…

As someone who's worked on similar situations, I don't expect any pat on the back, but see it as my responsibility to make sure it doesn't happen in the first place, and when it does, my responsibility to fix it without complaint. Some Atlassian customers might have had much more severe (mental health and other) problems than the Atlassian staff, so I think it can be perceived as mildly solipsistic to be praising the…

Speaking as someone who is no longer involved at the coal face of things like that, my role in an equivalent incident resolution would be to attend some of the 3-hourly calls (this frequency sounds extreme past the first 3-4days btw), and therefore having nothing to lose from siding with you on this, I’ll disagree. Accountability would be mine, responsibility only to fix of the team. I.e. the team needs many pats on the back, as they break their backs to solve something the organisation has done wrong, such as in this case not implementing proper controls on the processes. It’s never about the running code, it’s about how it ended up in production. People make mistakes, are assigned work outside their comfort zone, are junior. And these failure modes are also operational growth modalities. But the processes need to support that growth and focusing on the debugging of the issue is necessary for resolution but is wrong when the focus moves to long term avoidance.

You can praise staff but cannot praise the customer unfortunately as that would be inappropriate. You can work on relationship of course, by not limiting your reconciliation to contractual credits on SLA. These will mean nothing to the customer staff that took significant part of the hit in work stress. Not sure what would be a good way to regain the trust at that level.

Edit: fixed accountability vs responsibility in wrong order

Re: Post-incident review on the Atlassian April 2022 outage

#80

While the post-mortem is thorough, it misses key details on what companies experienced who were unlucky enough to be caught out by this outage. For example, it fails to mention how impacted customers lost access to certain Atlassian services for up to ~2 weeks: JIRA, Confluence, OpsGenie. But not others like Trello or BitBucket. Of these, losing access to OpsGenie for this long was a massive problem, dwarfing most ot…

It should be shouted from the rooftops: don't switch services just to save money if the result is potentially worse business outcomes. Why save a tiny bit of cash if it puts your business at risk?
Post reply on HN