Live data from Hacker News

Inside the longest Atlassian outage

newsletter.pragmaticengineer.com

661–670 of 772 posts

Re: Inside the longest Atlassian outage

#661
post #632

Earlier quoted context omitted.

People who call for other people's firings in organizations that they have no visibility into are so weird. This post reads like a tech outage's version of cancel culture where trying to find someone to blame and skewer for an injustice is more important than actually determining how much (if any) blame they deserve for it Also posting a LinkedIn event photos with people's real names and pictures in a top post on HN…

I agree. Customer success is a support role, this was an engineering mistake. Can't blame support for something an engineer did.

Seems like you're still assigning blame. Incidents are rarely if ever monocausal. The fantastic and accurate point the GP made is that fingerpointing is pointless. Much better to seek to learn and understand, which is always difficult but definitely can't be done from the sidelines without speaking to those involved.

Re: Inside the longest Atlassian outage

#662
post #632

Earlier quoted context omitted.

People who call for other people's firings in organizations that they have no visibility into are so weird. This post reads like a tech outage's version of cancel culture where trying to find someone to blame and skewer for an injustice is more important than actually determining how much (if any) blame they deserve for it Also posting a LinkedIn event photos with people's real names and pictures in a top post on HN…

How on earth did you rope cancel culture into this conversation?

Apparently getting fired for incompetence is now cancel culture lol. Sign me up for that world for sure!

Re: Inside the longest Atlassian outage

#663

I've repeatedly asked Atlassian if: 1. They can confirm that they have backups of our data (about a thousand stories, substantial confluence, opsgenie history, and three service desks). 2. Will our integrations, configuration, and customizations also be recovered, or will we need to rebuild those once our data is recovered? I have received no response, and no human is even willing to acknowledge those questions. The…

I was down, my instance is fully restored right now. 1. They do, every 25 hours, via snapshot. I have spoked to their team since the incident and that same thing is in this article. 2. Yes, they recover all of it. Some things have had issues, external mailboxes attached to service management projects, some attachment rendering slowness. Filters needing to be overlayed into our instance again, but otherwise it is runn…

Life saving Jira tickets?

Re: Inside the longest Atlassian outage

#664

Earlier quoted context omitted.

How on earth did you rope cancel culture into this conversation?

Apparently getting fired for incompetence is now cancel culture lol. Sign me up for that world for sure!

So the hypothetical UX designer for a new mobile app showed serious incompetence by allowing the infrastructure cloud team to mess up ops?

whether this is cancel culture not it definitely looks like finding an easy target to shout at

Re: Inside the longest Atlassian outage

#665
post #506

Earlier quoted context omitted.

Which is becoming more and more difficult due to them focusing on Cloud Products (my on-prem renewal jumped almost 8x this year). I’d rather use request tracker or bugzilla over Atlassian these days

Try Kitemaker, it’s a YComb backed JIRA alternative. https://kitemaker.co/

This landing page is exceedingly light on explaining what problem they solve for me.

Re: Inside the longest Atlassian outage

#666
post #411

This is extremely poor for a large SaaS company. A standard RFP question for SaaS should be: - Can you restore data for a single customer, and if so, what is the RTO for that operation? A smaller SaaS could be excused for only thinking about full database restores. When you're a scrappy upstart, thinking about hypotheticals is less important than survival. But for any decent size multi-tenanted SaaS, it's imperative…

> Atlassian - MUST DO BETTER. it’s not like people will stop using jira and confluence, lol they basically have a monopoly there

Nobody likes using this stuff though. I think the new github boards might give jira a run for their money given the cost... free.

Re: Inside the longest Atlassian outage

#667
post #632

Engineering mistakes happen. The most inexcusable thing is not communicating with the paying customers who have been affected for over a week. Atlassian's Global Head of Customer Success probably should have been fired but here she is promoting Atlassian Cloud on LinkedIn three days ago: https://www.linkedin.com/mwlite/in/gertie-rizzo-5b70061 Actually reading a bit more, it seems like their customer team was partying…

People who call for other people's firings in organizations that they have no visibility into are so weird. This post reads like a tech outage's version of cancel culture where trying to find someone to blame and skewer for an injustice is more important than actually determining how much (if any) blame they deserve for it Also posting a LinkedIn event photos with people's real names and pictures in a top post on HN…

The original commenter (OC) did not call for anyone to be fired. They stated an opinion; they did not make a demand.

The OC is directing us to information made public by Atlassian team members themselves. The OC did not make the information public. Anyone could find that information with little effort.

I’m not really sure why you attribute such intensity to the OC’s rather benign comments.

Re: Inside the longest Atlassian outage

#668
post #571

Earlier quoted context omitted.

You seem to be implying that customer success is a customer support role. It isn't. Customer success is about helping customers get the greatest business value from your product. They are not going to be munging databases and wrangling backups. Support engineers (among others) do that work.

Well right now customers are getting zero value. It’s a bad look to be partying during an all hands on deck emergency.

More hands on that deck probably don't help.

Re: Inside the longest Atlassian outage

#669

Wow would you look at that, a complete Atlassian puff piece got published in the WSJ just hours ago. How peculiar that the biggest active outage in the history of this company is not mentioned once in this "article". I'm left to assume that PR teams can plant whatever they see fit in the WSJ at a moment's notice. I guess that's what passes for journalism these days. https://www.wsj.com/articles/atlassian-puts-easy-to…

[deleted]

Re: Inside the longest Atlassian outage

#670

Earlier quoted context omitted.

I am Dev Ops, and not Ops. So I try to not waste time with self hosting as much as possible.

DevOps really is not just doing DevOps on cloud platforms and SaaS. Besides, the sysadmin aspects of self hosting should be handled by, well, sysadmin. DevOps should be handling other aspects like developing solutuons necessary to have things (in this case Jira) work together with other systems. (Among other responsibilities) Though DevOps can implemented in different ways with responsibilities that are different fro…

Yes, I should have written a bit more. It is definitely a preference. Currently, in my team Devops is a one man show (me). So I have to be careful with how much burden I allow. There are things that make sense to self host, I agree. I am just trying to avoid becoming classical IT and having to provide lots of end user support and such.
Post reply on HN