You missed your recovery time objective by ~2 weeks and did not properly communicate the issue from senior management until around a week into the outage. It is great to hear about how the company plans to do better, let's see how the next outage improves things.
Post-incident review on the Atlassian April 2022 outage
111–120 of 157 posts
Re: Post-incident review on the Atlassian April 2022 outage
#112Earlier quoted context omitted.
But you don’t need a paging system if your services don’t go down. Isn’t the fact you can’t deal without one for two weeks an indictment of your own practices?
Sure and you don't need emergency locator transmitters if your aircraft doesn't crash. When you're ready to prove that your services "don't go down" send me an email and I'll come work for you.
Re: Post-incident review on the Atlassian April 2022 outage
#113I can only imagine the way the person who pulled the trigger on the deletion script felt the moment they realized what had happened. I’ve been there with much less significant incidents when a “routine” change turned into a potentially resume generating event. It’s not fun. Ultimately the responsibility is with the organization that made it possible for an event of that scale to happen rather than the individual pers…
Thomas J. Watson famously said: Recently, I was asked if I was going to fire an employee who made a mistake that cost the company $600,000. No, I replied, I just spent $600,000 training him. Why would I want somebody to hire his experience?
Re: Post-incident review on the Atlassian April 2022 outage
#114Earlier quoted context omitted.
Thomas J. Watson famously said: Recently, I was asked if I was going to fire an employee who made a mistake that cost the company $600,000. No, I replied, I just spent $600,000 training him. Why would I want somebody to hire his experience?
This quip never made much sense. Mistakes almost certainly follow some kind of pareto distribution instead of it being evenly distributed. That's why 2% of doctors are responsible for 39% of malpractice lawsuits. If your doctor cut off the wrong leg and you sued (and won) $2 million, would go back to the same doctor and just chalk it up as "well, now he's got $2 million of training"?
I see it as: Your employee just learned a valuable lesson, and you definitely don't want to hire another new employee to make that same mistake again.
Ultimately, it's the owner taking responsibility for the fuck up because they hired the employee, but still standing behind their decision to hire the employee.
If your employee is an idiot and makes up 39% of mistakes, of course you won't keep them around when they make a big one. But most employees are not fuck-ups, as you alluded to with 2% of doctors making a majority of malpractice suits.
Re: Post-incident review on the Atlassian April 2022 outage
#115While the post-mortem is thorough, it misses key details on what companies experienced who were unlucky enough to be caught out by this outage. For example, it fails to mention how impacted customers lost access to certain Atlassian services for up to ~2 weeks: JIRA, Confluence, OpsGenie. But not others like Trello or BitBucket. Of these, losing access to OpsGenie for this long was a massive problem, dwarfing most ot…
But you don’t need a paging system if your services don’t go down. Isn’t the fact you can’t deal without one for two weeks an indictment of your own practices?
Re: Post-incident review on the Atlassian April 2022 outage
#116Earlier quoted context omitted.
"what about"-isms don't remove the experiences of people with jobs less stressful than rail clearing specialists.
I'm curious about your take on the cheapening or dilution of words in language nowadays. One isnt hurt but traumatized, one isn't "once bitten twice shy" but suffering from PTSD. It's assault if one gets in a fight and so on. I sometimes feel we're running out of language to distinguish the truly extreme from the quotidian. Same thing with how GP phrased it. A hard week of work comes across as having gone to battle a…
Any argument that uses this qualification should be accompanied with data.
You haven't specified anything, not even a rough time frame. Your statement is so vague nearly anything can be projected onto it.
Also a sidenote: if you're going to make an argument about language and definitions it seems like a good idea to actually know what PTSD is before using it as an example.
Re: Post-incident review on the Atlassian April 2022 outage
#117Earlier quoted context omitted.
It should be shouted from the rooftops: don't switch services just to save money if the result is potentially worse business outcomes. Why save a tiny bit of cash if it puts your business at risk?
The saying "Cheap is expensive" comes to mind
This is not very far away from AWS hosting their own status pages.
Re: Post-incident review on the Atlassian April 2022 outage
#118I sincerely hope all of the people who worked on the recovery effort are okay, and are being well supported and strongly encouraged to tend their mental health. I have no personal investment in Atlassian products—if anything, unrelated to this or any incident, I could happily never use them again. But the people who work there are human, and I know what kind of a toll a protracted recovery effort can take. I know it…
As someone who's worked on similar situations, I don't expect any pat on the back, but see it as my responsibility to make sure it doesn't happen in the first place, and when it does, my responsibility to fix it without complaint. Some Atlassian customers might have had much more severe (mental health and other) problems than the Atlassian staff, so I think it can be perceived as mildly solipsistic to be praising the…
Re: Post-incident review on the Atlassian April 2022 outage
#119Earlier quoted context omitted.
Is that real? That is so bad. Queue up next “Shall not disclose or notice our incompetence” clause.
https://www.atlassian.com/legal/cloud-terms-of-service > 3.3. Restrictions. Except as otherwise expressly permitted in these Terms, you will not: > (i) publicly disseminate information regarding the performance of the Cloud Products; or (j) encourage or assist any third party to do any of the foregoing.
Re: Post-incident review on the Atlassian April 2022 outage
#120I can't say that I've ever been a fan of Atlassian or their products, but this blog post makes it sounds like they've at least learned the right lessons from this: 1. Establish universal "soft deletes" across all systems. 2. Better DR for multi-site, multi-product incidents. 3. Fix their incident-management process for large-scale incidents. 4. Fix their incident communications. Regarding #4 in particular: "Rather th…
I tend to agree. However, there's one point that makes me skeptical: there are no organizational changes, or changes to leadership, or anything in that direction. This sounds like "the tech guys screwed up, culture and management is fine here". Which it might be, or it might not. I would have loved to see 5. We will stop pushing customers so hard towards using our cloud for one, but that wouldn't be convenient for At…