Live data from Hacker News

Post-incident review on the Atlassian April 2022 outage

atlassian.com

111–120 of 157 posts

Re: Post-incident review on the Atlassian April 2022 outage

#111
> During this incident, we missed our RTO but met our RPO.

You missed your recovery time objective by ~2 weeks and did not properly communicate the issue from senior management until around a week into the outage. It is great to hear about how the company plans to do better, let's see how the next outage improves things.

Re: Post-incident review on the Atlassian April 2022 outage

#112
post #103
post #90

Earlier quoted context omitted.

But you don’t need a paging system if your services don’t go down. Isn’t the fact you can’t deal without one for two weeks an indictment of your own practices?

Sure and you don't need emergency locator transmitters if your aircraft doesn't crash. When you're ready to prove that your services "don't go down" send me an email and I'll come work for you.

Even if they do, I wouldn’t want to work for them

Re: Post-incident review on the Atlassian April 2022 outage

#113

I can only imagine the way the person who pulled the trigger on the deletion script felt the moment they realized what had happened. I’ve been there with much less significant incidents when a “routine” change turned into a potentially resume generating event. It’s not fun. Ultimately the responsibility is with the organization that made it possible for an event of that scale to happen rather than the individual pers…

Thomas J. Watson famously said: Recently, I was asked if I was going to fire an employee who made a mistake that cost the company $600,000. No, I replied, I just spent $600,000 training him. Why would I want somebody to hire his experience?

Lol do you should keep an idiot who shot a (looks at notes for what’s as cheap as 600k in the army) a javelin missile while drunk? It’s prudent to apply the concept of nuance. If some intern deleted all these accounts you should fire the guy responsible for letting things free enough for the intern to do it but you should fire them. Whatever lesson learned could be learned from a wiki page detailing the incident.

Re: Post-incident review on the Atlassian April 2022 outage

#114

Earlier quoted context omitted.

Thomas J. Watson famously said: Recently, I was asked if I was going to fire an employee who made a mistake that cost the company $600,000. No, I replied, I just spent $600,000 training him. Why would I want somebody to hire his experience?

This quip never made much sense. Mistakes almost certainly follow some kind of pareto distribution instead of it being evenly distributed. That's why 2% of doctors are responsible for 39% of malpractice lawsuits. If your doctor cut off the wrong leg and you sued (and won) $2 million, would go back to the same doctor and just chalk it up as "well, now he's got $2 million of training"?

> This quip never made much sense.

I see it as: Your employee just learned a valuable lesson, and you definitely don't want to hire another new employee to make that same mistake again.

Ultimately, it's the owner taking responsibility for the fuck up because they hired the employee, but still standing behind their decision to hire the employee.

If your employee is an idiot and makes up 39% of mistakes, of course you won't keep them around when they make a big one. But most employees are not fuck-ups, as you alluded to with 2% of doctors making a majority of malpractice suits.

Re: Post-incident review on the Atlassian April 2022 outage

#115
post #90

While the post-mortem is thorough, it misses key details on what companies experienced who were unlucky enough to be caught out by this outage. For example, it fails to mention how impacted customers lost access to certain Atlassian services for up to ~2 weeks: JIRA, Confluence, OpsGenie. But not others like Trello or BitBucket. Of these, losing access to OpsGenie for this long was a massive problem, dwarfing most ot…

But you don’t need a paging system if your services don’t go down. Isn’t the fact you can’t deal without one for two weeks an indictment of your own practices?

Well said - you don't need to write tests for your code if you don't write bugs!

Re: Post-incident review on the Atlassian April 2022 outage

#116

Earlier quoted context omitted.

"what about"-isms don't remove the experiences of people with jobs less stressful than rail clearing specialists.

I'm curious about your take on the cheapening or dilution of words in language nowadays. One isnt hurt but traumatized, one isn't "once bitten twice shy" but suffering from PTSD. It's assault if one gets in a fight and so on. I sometimes feel we're running out of language to distinguish the truly extreme from the quotidian. Same thing with how GP phrased it. A hard week of work comes across as having gone to battle a…

> nowadays

Any argument that uses this qualification should be accompanied with data.

You haven't specified anything, not even a rough time frame. Your statement is so vague nearly anything can be projected onto it.

Also a sidenote: if you're going to make an argument about language and definitions it seems like a good idea to actually know what PTSD is before using it as an example.

Re: Post-incident review on the Atlassian April 2022 outage

#117
post #83

Earlier quoted context omitted.

It should be shouted from the rooftops: don't switch services just to save money if the result is potentially worse business outcomes. Why save a tiny bit of cash if it puts your business at risk?

The saying "Cheap is expensive" comes to mind

Also "don't put all your eggs in one basket"

This is not very far away from AWS hosting their own status pages.

Re: Post-incident review on the Atlassian April 2022 outage

#118
post #70

I sincerely hope all of the people who worked on the recovery effort are okay, and are being well supported and strongly encouraged to tend their mental health. I have no personal investment in Atlassian products—if anything, unrelated to this or any incident, I could happily never use them again. But the people who work there are human, and I know what kind of a toll a protracted recovery effort can take. I know it…

As someone who's worked on similar situations, I don't expect any pat on the back, but see it as my responsibility to make sure it doesn't happen in the first place, and when it does, my responsibility to fix it without complaint. Some Atlassian customers might have had much more severe (mental health and other) problems than the Atlassian staff, so I think it can be perceived as mildly solipsistic to be praising the…

I think that’s a false trade off, you can empathise with both employees and customers at the same time. And while sure maybe a lot of the Atlassians might have fixed the issue “without complaint” as you said, it can still be a nightly stressful situation (probably even leading to burnout) and I hope they are looked after even if they don’t speak up.

Re: Post-incident review on the Atlassian April 2022 outage

#119

Earlier quoted context omitted.

Is that real? That is so bad. Queue up next “Shall not disclose or notice our incompetence” clause.

https://www.atlassian.com/legal/cloud-terms-of-service > 3.3. Restrictions. Except as otherwise expressly permitted in these Terms, you will not: > (i) publicly disseminate information regarding the performance of the Cloud Products; or (j) encourage or assist any third party to do any of the foregoing.

That should tell you everything you need to know about atlassian as a company and about how good their cloud product is.

Re: Post-incident review on the Atlassian April 2022 outage

#120

I can't say that I've ever been a fan of Atlassian or their products, but this blog post makes it sounds like they've at least learned the right lessons from this: 1. Establish universal "soft deletes" across all systems. 2. Better DR for multi-site, multi-product incidents. 3. Fix their incident-management process for large-scale incidents. 4. Fix their incident communications. Regarding #4 in particular: "Rather th…

I tend to agree. However, there's one point that makes me skeptical: there are no organizational changes, or changes to leadership, or anything in that direction. This sounds like "the tech guys screwed up, culture and management is fine here". Which it might be, or it might not. I would have loved to see 5. We will stop pushing customers so hard towards using our cloud for one, but that wouldn't be convenient for At…

By keeping the management you have a chance they learn a lot from that unique experience. Changing leadership would be the PR move.
Post reply on HN