Live data from Hacker News

Inside the longest Atlassian outage

newsletter.pragmaticengineer.com

631–640 of 772 posts

Re: Inside the longest Atlassian outage

#631

I've repeatedly asked Atlassian if: 1. They can confirm that they have backups of our data (about a thousand stories, substantial confluence, opsgenie history, and three service desks). 2. Will our integrations, configuration, and customizations also be recovered, or will we need to rebuild those once our data is recovered? I have received no response, and no human is even willing to acknowledge those questions. The…

I was down, my instance is fully restored right now.

1. They do, every 25 hours, via snapshot. I have spoked to their team since the incident and that same thing is in this article. 2. Yes, they recover all of it. Some things have had issues, external mailboxes attached to service management projects, some attachment rendering slowness. Filters needing to be overlayed into our instance again, but otherwise it is running fine again.

Not sure what to tell you other than they are fixing life saving companies first, then the rest. That is what they have told us.

Re: Inside the longest Atlassian outage

#632

Engineering mistakes happen. The most inexcusable thing is not communicating with the paying customers who have been affected for over a week. Atlassian's Global Head of Customer Success probably should have been fired but here she is promoting Atlassian Cloud on LinkedIn three days ago: https://www.linkedin.com/mwlite/in/gertie-rizzo-5b70061 Actually reading a bit more, it seems like their customer team was partying…

People who call for other people's firings in organizations that they have no visibility into are so weird. This post reads like a tech outage's version of cancel culture where trying to find someone to blame and skewer for an injustice is more important than actually determining how much (if any) blame they deserve for it

Also posting a LinkedIn event photos with people's real names and pictures in a top post on HN along with provocative framing like "partying in Vegas" about this outage is pretty shitty. Atlassian is a big company and has a huge engineering and sales department, you have no idea if any of those people at the Vegas event had anything to contribute with the outage response. For example in your link there are literally people talking about some new mobile app their team is launching at the event, I doubt any of those people are involved in the outage response.

No one posting on HN unless they work at Atlassian in a leadership role is in any position to even start assigning blame, call for firings or publicly shaming people (the later of which you shouldn't be doing even if you 100% knew for a fact that they were at fault).

Do better.

Re: Inside the longest Atlassian outage

#633
post #568

Earlier quoted context omitted.

Hi, I'm Mike and I work in Engineering at Atlassian. Here's our approach to backup and data management: https://www.atlassian.com/trust/security/data-management - we certainly have the backups and have a restore process that we keep to. However, this incident stressed our ability to do this at scale, which has led to the very long times to restore.

400 tenants doesn't seem like that much scale though...? What will happen if there's an incident affecting more than 0.18% of tenants?

It's 400 tenants scattered across all their servers. So they are most likely having to build out servers to pull the data then put it in place. 10x the problem that just restoring a single server would be.

Re: Inside the longest Atlassian outage

#634

Regarding the backup restores: I once worked a company that had a data loss issue. There was nothing else we could do, we had exhausted every option we had over almost 40 hours. At the end of the second day, it was decided to restore from backup. We had done this before, as a test. It took about 12 hours to restore the data and another 12 hours to import the data and get back up and running. One small thing was diffe…

Hi, I'm Mike and I work in Engineering at Atlassian. Here's our approach to backup and data management: https://www.atlassian.com/trust/security/data-management - we certainly have the backups and have a restore process that we keep to. However, this incident stressed our ability to do this at scale, which has led to the very long times to restore.

I think this document and incident is a decent example of common DR planning failure patterns.

It is explained here that Atlassian runs regular DR planning meetings with the engineers spending time planing out potential scenarios, as well as quarterly tests of backups and tracking findings from them.

So, with those two things happening, I the imagine recovery time objectives of That doesn't even come close to the recovery time we are currently seeing now however. We're coming up on 2 orders of magnitude more than that.

The above doc seems pretty far our of line with what is currently happening.

Re: Inside the longest Atlassian outage

#635
post #603

Earlier quoted context omitted.

Why should sales take the blame when it's engineering's problem? If Sales promised a feature to a customer that was infeasible _then_ it should be Sales problem, but engineers made a mistake so only engineers can clean up the mistake.

You’re missing something huge. Sales almost certainly sold features that meant engineering couldn’t focus on things the engineers thought were valuable, like better automation.

Leadership is responsible for deciding which of those to prioritize.

Re: Inside the longest Atlassian outage

#636

Earlier quoted context omitted.

25 years ago the clutch in my beater truck was slipping. I was 16 years old, making $50 a week and had very little in savings. I took that truck to a shop within walking distance of my job. 2 hours later I walked back to see what they found. I figured it would be several hundred dollars for a new clutch, and I'd have to borrow money or something to get it done. I talked to the owner who told be it was an adjustment o…

Your anecdote is nice, and sure it can be good advertising to give stuff away for free, But it doesn't really apply here. If you were charged $123.98 and you said, "hey, I told you where the problem was, why am I being charged a diagnostics and driving fee?" and they corrected it by telling you the whole thing is on the house, is that not good business sense? Even by your own admission, you would have gladly paid tha…

> If you were charged $123.98 and you said, "hey, I told you where the problem was, why am I being charged a diagnostics and driving fee?" and they corrected it by telling you the whole thing is on the house, is that not good business sense?

No. I'll be happy that I saved on the money, but I won't trust them in the future. They're now "the place that tries to get away with things" in my mental Rolodex. Better to stick with the fee and know their value. (I didn't tell them where the problem was. All I knew was that the clutch wasn't grabbing anymore. I assumed it needed a whole new clutch.)

> Even by your own admission, you would have gladly paid that $123.98 with no issues and you wouldn't have been mad about it. So from a business perspective, if they can provide a service, get paid for it, and the customer has no qualms or issues with the transaction whatsoever, in what way is that hurting the brand or being cynical?

It would have been a fine decision, sure. But in that case that would likely have been the only business I did with them. Not out of spite or anger, but because I'd have no reason to pick them for future business. I would instead ask friends for recommendations, or pick some place closer to my future residences.

But what actually happened was that I was the one steering people to them. I also went out of my way to return to them for brake jobs, simple oil changes, etc. I was a loyal customer, and probably spent or caused others to spend over $5,000 there.

He had absolutely know way of knowing that would result. But if you just treat people right, the way you'd want them to treat you, you build a reputation. It pays back.

I know this story comes off a bit pollyanna. I get it. For a cynical and non-altruistic explanation: when it takes a technician literally 5 minutes to twist an adjustment nut and verify that was all there was to it, stop and think about the bigger opportunity before you robotically mark '1.00' in the "LBR HRS" field on an invoice. Especially if you're operating in a field that's notorious for rip offs.

> I think that's a much more business-wise action to take than to give away your services.

I'm not saying businesses should give away major services. But they should avoid the temptation to nickel-and-dime as well. That's on the other end of the optimization curve. Not good business.

Re: Inside the longest Atlassian outage

#637

What's a good Jira replacement? Redmine? Phabricator? OpenProject? Just leaving the jira server alone and hoping there's no new and exciting zero-days? One thing is clear, these guys are a bunch of cowboys who can't be trusted with any amount of data.

I'm not at all familiar but a tweet linked from the OP and written by the author plugs https://linear.app/

Linear has offered free services to users impacted by Atlassian's outage through the end of the year. I took a look at it (we aren't impacted), and notice it can import tickets from Jira, and also has a "Jira Link" where you can use Linear as a kind of front-end to Jira if you aren't ready to go all in on Jira.

When we chose Jira, one of the points that was made was: If we decide to leave Jira, there will almost certainly be an importer from Jira to the new system. Which does seem to be true. We came to Jira from Fogbugz, and I spent the better part of a month writing tools to import our tickets and wikis. Jira had a Fogbugz importer, but it was horribly broken.

Looking at Linear, there is no such escape hatch, or indeed, searching the docs I saw no "export" or "backup" capability at all.

Re: Inside the longest Atlassian outage

#638
post #506

Earlier quoted context omitted.

Rather, customers must stop using Atlassian cloud services.

Which is becoming more and more difficult due to them focusing on Cloud Products (my on-prem renewal jumped almost 8x this year). I’d rather use request tracker or bugzilla over Atlassian these days

Try Kitemaker, it’s a YComb backed JIRA alternative.

https://kitemaker.co/

Re: Inside the longest Atlassian outage

#639
Honest question here: The companies impacted by this, are they not taking backups of their Jira/Confluence/Bitbucket instances? Or is this outage impacting the ability to import those backups?

There are some Python scripts that will back up Jira and Confluence. I whipped up a quick script that gets a list of all our bitbucket repos and then it clones those daily as well.

Re: Inside the longest Atlassian outage

#640

Earlier quoted context omitted.

Or some overzealous engineer said hey guys let's delete all data 7 days after an account is canceled. This is called over optimizing.

Such a decision is just as likely to have come from the legal/compliance team as an engineer. Data you no longer have clear consent or a legitimate business need to store is a liability, and if you operate in Europe, potentially illegal to continue storing.

It’s amazing how much stupid shit we do to keep the legal guys happy while their bosses are busy engaging in tax evasion, graft, bribery, fraud, embezzlement, illegal dumping, sexual harassment, sexual assault, statutory rape, solicitation to commit murder, and my personal favorite and I’m sure yours too: human trafficking.

But sure, we can break all of our users to avoid the possibility of you having to write some legal briefs and us paying a small fine for keeping data 7 days instead of three.

Post reply on HN