Live data from Hacker News

Post-incident review on the Atlassian April 2022 outage

atlassian.com

91–100 of 157 posts

Re: Post-incident review on the Atlassian April 2022 outage

#91
post #70

I sincerely hope all of the people who worked on the recovery effort are okay, and are being well supported and strongly encouraged to tend their mental health. I have no personal investment in Atlassian products—if anything, unrelated to this or any incident, I could happily never use them again. But the people who work there are human, and I know what kind of a toll a protracted recovery effort can take. I know it…

As someone who's worked on similar situations, I don't expect any pat on the back, but see it as my responsibility to make sure it doesn't happen in the first place, and when it does, my responsibility to fix it without complaint. Some Atlassian customers might have had much more severe (mental health and other) problems than the Atlassian staff, so I think it can be perceived as mildly solipsistic to be praising the…

I find it hard to imagine mental health being affected negatively by atlassian services disappearing.

Re: Post-incident review on the Atlassian April 2022 outage

#92

Earlier quoted context omitted.

Note that the ToS also forbid Cloud users from “disseminating information” about “the performance of the products”. So you can’t say it’s slow or unperforming.

I thought, in terms of Jira and Confluence at least, it was just accepted that being slow and underperforming was the status quo and if it was running at a speed you'd consider normal, that's an exception (and cause for alarm... like "did that actually save or is there a silent JS error not being displayed?").

Sure, it’d be extremely hard for their cloud offering to be slower than our on-premise installation.

Re: Post-incident review on the Atlassian April 2022 outage

#93

Implementing soft deletes is a lessons every developer learns early in their career. The fact that Atlassian did not implement that in their cloud, is mind boggling. Great case study of a monumental fuck up!

So many opportunities missed to avoid this! Look at one ID and ensure it is what you expect it to be. Run the script in a dryrun mode. Run the script for 1 customer. Probably more!

And don’t make a “universal delete” script in the first place…

Re: Post-incident review on the Atlassian April 2022 outage

#94

Earlier quoted context omitted.

Quoted post unavailable.

"what about"-isms don't remove the experiences of people with jobs less stressful than rail clearing specialists.

I'm curious about your take on the cheapening or dilution of words in language nowadays. One isnt hurt but traumatized, one isn't "once bitten twice shy" but suffering from PTSD. It's assault if one gets in a fight and so on. I sometimes feel we're running out of language to distinguish the truly extreme from the quotidian. Same thing with how GP phrased it. A hard week of work comes across as having gone to battle and being left with deep wounds.

Still, I do mostly get your point.

Re: Post-incident review on the Atlassian April 2022 outage

#95

While the post-mortem is thorough, it misses key details on what companies experienced who were unlucky enough to be caught out by this outage. For example, it fails to mention how impacted customers lost access to certain Atlassian services for up to ~2 weeks: JIRA, Confluence, OpsGenie. But not others like Trello or BitBucket. Of these, losing access to OpsGenie for this long was a massive problem, dwarfing most ot…

As a customer I would not buy or use an Atlassian product in a 1000 years.

14 days without pagers ...

From friends and people I know I heard nothing good about Atlassian products. And that was before that 2 weeks downtime.

It looks like the products are duct together with duct tape, spit and a little bit of dirt.

Re: Post-incident review on the Atlassian April 2022 outage

#96
post #30
post #24

I am very curious if they used a Jira board during this crisis for issue tracking. Because then they would have more than 4 lessons learned.

What you're basically suggesting is that feature development at Atlassian moves at such a glacial speed because of course they're using Jira to manage it. This kind of blows my mind right now.

I chuckled hard hahaha

Re: Post-incident review on the Atlassian April 2022 outage

#97

I can only imagine the way the person who pulled the trigger on the deletion script felt the moment they realized what had happened. I’ve been there with much less significant incidents when a “routine” change turned into a potentially resume generating event. It’s not fun. Ultimately the responsibility is with the organization that made it possible for an event of that scale to happen rather than the individual pers…

I was on call “back in the day” when a customer rang to tell me that she’d meant to type: # find /tmp -exec rm {} /; But instead she’d typed # find / tmp -exec rm {} /; I learned a lot about Unix (and, as it happens, X windows) file systems that week. Note that this was long before tools like sudo existed, where cleaning up tmp was often a manual process. There is no way we could blame a person for a single keystroke…

Good old days deleting an entire server ... learned that the hard way.

Re: Post-incident review on the Atlassian April 2022 outage

#98

I sincerely hope all of the people who worked on the recovery effort are okay, and are being well supported and strongly encouraged to tend their mental health. I have no personal investment in Atlassian products—if anything, unrelated to this or any incident, I could happily never use them again. But the people who work there are human, and I know what kind of a toll a protracted recovery effort can take. I know it…

Quoted post unavailable.

Okay you’re actually just being a jerk. I’ll take my two years of chronic pain and go live for that excitement. But you can go take your gob’s sakes and shove them.

Re: Post-incident review on the Atlassian April 2022 outage

#99

Earlier quoted context omitted.

"what about"-isms don't remove the experiences of people with jobs less stressful than rail clearing specialists.

I'm curious about your take on the cheapening or dilution of words in language nowadays. One isnt hurt but traumatized, one isn't "once bitten twice shy" but suffering from PTSD. It's assault if one gets in a fight and so on. I sometimes feel we're running out of language to distinguish the truly extreme from the quotidian. Same thing with how GP phrased it. A hard week of work comes across as having gone to battle a…

You’re projecting a whole lot of language that’s not in use here so maybe you should consider that.

Re: Post-incident review on the Atlassian April 2022 outage

#100

I can only imagine the way the person who pulled the trigger on the deletion script felt the moment they realized what had happened. I’ve been there with much less significant incidents when a “routine” change turned into a potentially resume generating event. It’s not fun. Ultimately the responsibility is with the organization that made it possible for an event of that scale to happen rather than the individual pers…

Thomas J. Watson famously said: Recently, I was asked if I was going to fire an employee who made a mistake that cost the company $600,000. No, I replied, I just spent $600,000 training him. Why would I want somebody to hire his experience?

This quip never made much sense.

Mistakes almost certainly follow some kind of pareto distribution instead of it being evenly distributed.

That's why 2% of doctors are responsible for 39% of malpractice lawsuits. If your doctor cut off the wrong leg and you sued (and won) $2 million, would go back to the same doctor and just chalk it up as "well, now he's got $2 million of training"?

Post reply on HN