Live data from Hacker News

Inside the longest Atlassian outage

newsletter.pragmaticengineer.com

141–150 of 772 posts

Re: Inside the longest Atlassian outage

#141

A few years ago we didn't renew our subscription on time because we got the email over Christmas break, and iirc they deleted all of our data in less than two weeks. They were eventually able to manually restore it from backups, but they restored it incorrectly so there was a bunch of stuff broken. This whole thing isn't even remotely surprising to me.

Did you continue as their customer after that?

Nope. I exported our data after they restored the backup and then we cancelled less than a month later. Like I obviously understand suspending our logins, but why would you ever delete someone's data when it's literally only 160 KB of text? The whole thing made zero sense.

Re: Inside the longest Atlassian outage

#142
We use on-premises setups for almost everything (we generally avoid cloud solutions to have full control of our data), sometimes (approximately once a month) it goes down for a few minutes which already feels like a torture because all our processes depend on it, I can't imagine having no access to it for several weeks, all our work would stop to a halt... The office of the guy who administers on-premise servers is literally next door, all it takes is to make a visit to him and everything works again after 5 minutes. Reading horror stories like this (Slack being down, Atlassian being down, no one knows what is happening and when it will end etc.), I wonder why many companies choose cloud solutions for critical business processes. Is it pricing? Ease of use? I can understand why very small companies would choose it, but I don't understand why a medium/large business would choose anything but an on-premises setup.

Re: Inside the longest Atlassian outage

#143

> it takes between 4 and 5 elapsed days to hand a site back to a customer. Atlassian's SLA page says, Premium Cloud Products 99.9% That's 43 minutes of downtime per month. That works out to, Atlassian can't have any more downtime for the next 14 years. Are SLAs even real? I'm being slightly facetious. From the page text it's just a threshold after which I think you're entitled to some money back for that month.

SLAs aren't real unless there's a contractual consequence for not meeting them. And a couple of percent discount on services for the extra downtime isn't really a meaningful consequence.

I was just thinking that there's a hysteresis function here: the service is worth much more to your team after you've wired your whole process into it than before you joined.

Offering you a free month or whatever doesn't acknowledge all the person-hours lost.

Re: Inside the longest Atlassian outage

#144
post #58

Earlier quoted context omitted.

The answer is medium to large companies. Jira is a tool that can satisfy hundreds of different teams’ work management needs without having to buy dozens of different products. The fact that it’s so feature packed and customizable is the point. I think the complainers are not really investing the time in to change project settings to fit their needs. My only complaint about the Atlassian suite is the performance of Ji…

How do you change the markup language to be consistent between Jira and Confluence? How do you eliminate all non-task ticket types in a Jira board and allow any ticket to be a child of any other ticket? It’s hard to configure away complexity from a product if it’s designed to be complicated.

Re 1: I'm not sure why that's a necessity beyond a notion of consistency. I find that major wiki editors are not often major ticket creators, and these are different products with different audiences at the end of the day. Also, Confluence uses a WYSIWYG editor, so it's rare to need to think about the markup.

Re 2: Set the project's issue type scheme to one that only allows tasks and subtasks. That gets you one level of nesting. (And even though task and subtasks are different issue types, changing from one to the other is trivial since they have identical fields.) Allowing epics gets you another at the top level. That's a bit limited, but wouldn't arbitrary nesting be even more complex?

Re: Inside the longest Atlassian outage

#145
Regarding the backup restores:

I once worked a company that had a data loss issue. There was nothing else we could do, we had exhausted every option we had over almost 40 hours. At the end of the second day, it was decided to restore from backup.

We had done this before, as a test. It took about 12 hours to restore the data and another 12 hours to import the data and get back up and running.

One small thing was different this time, and it had huge consequences. As a cost-saving measure, an engineer had changed the location of our backups to the cold-storage tier offered by our cloud provider. All backups, not just 'old' ones.

This added 2 additional days to our recovery time, for a total of five days. Interestingly enough, even though we offered a full month's refund to all of our customers, not even half of them took us up on it.

Re: Inside the longest Atlassian outage

#146

Earlier quoted context omitted.

Not being able to selectively restore data for a subset of users might also indicate that users are not fully isolated from each other, which is worrying for technical and nontechnical reasons.

There is nothing non-technical that matters. If we start acting like it does, we incredibly poor decisions that in fact have nothing to do with physical reality, and quickly arrive at unworkable technology.

Non-technical reasons include "legal" and "compliance", which often matters a fair bit. I am not disagreeing that non-technical requirements occasionally lead to poor decisions, for some value of poor.

Re: Inside the longest Atlassian outage

#148
post #142

We use on-premises setups for almost everything (we generally avoid cloud solutions to have full control of our data), sometimes (approximately once a month) it goes down for a few minutes which already feels like a torture because all our processes depend on it, I can't imagine having no access to it for several weeks, all our work would stop to a halt... The office of the guy who administers on-premise servers is l…

Cloud solutions can work well. I've used GitHub, Azure Devops, and BitBucket (another wonderful atlassian product /s) and BitBucket frequently craps out, multiple times a week. We need to rerun builds in TeamCity because BitBucket stops talking to it.

Re: Inside the longest Atlassian outage

#150
post #142

We use on-premises setups for almost everything (we generally avoid cloud solutions to have full control of our data), sometimes (approximately once a month) it goes down for a few minutes which already feels like a torture because all our processes depend on it, I can't imagine having no access to it for several weeks, all our work would stop to a halt... The office of the guy who administers on-premise servers is l…

You're assuming every team would have better uptime with in-house solutions

I think many would have worse uptime even with more headcount

Post reply on HN