Live data from Hacker News

Inside the longest Atlassian outage

newsletter.pragmaticengineer.com

731–740 of 772 posts

Re: Inside the longest Atlassian outage

#731
post #632

Earlier quoted context omitted.

People who call for other people's firings in organizations that they have no visibility into are so weird. This post reads like a tech outage's version of cancel culture where trying to find someone to blame and skewer for an injustice is more important than actually determining how much (if any) blame they deserve for it Also posting a LinkedIn event photos with people's real names and pictures in a top post on HN…

The main article is about how Atlassian have, out of nowhere, ceased business operations for 400 of their customers. They are no longer a going concern and this could happen to anyone who uses their products which is (alas) quite a lot of us. That’s my perspective. Reading your comment it feels like between the lines you don’t think this is as serious as other people. If you have a different perspective, could you co…

Do you realize just how disconnected sales and operations are? Look at your organization chart. They connect at the CEO.

My previous company lost customer data. Someone deleted a previous employee's account which surprise contained a production customer website. We didn't halt sales and outreach.

I mean RSA lost all of their encryption seeds in 2011. Every account was compromised. They still do sales today.

Re: Inside the longest Atlassian outage

#732
post #658

Earlier quoted context omitted.

Linear has offered free services to users impacted by Atlassian's outage through the end of the year. I took a look at it (we aren't impacted), and notice it can import tickets from Jira, and also has a "Jira Link" where you can use Linear as a kind of front-end to Jira if you aren't ready to go all in on Jira. When we chose Jira, one of the points that was made was: If we decide to leave Jira, there will almost cert…

It's in the settings, "Import / Export" in the sidebar. We support CSV export or alternatively you can use the GraphQL API. You're right we need add a page in the docs. It was only mentioned in few pages in passing. For now I added a section here until we can make a full page about it: https://linear.app/docs/workspaces#export-workspace-data If you are looking anything else hit cmd+k to search the docs.

Thanks for the reply. Linear looks pretty slick, I'll probably give it a try with the Jira Link, and get some experience with it without having to do a whole conversion plus get buyin from the rest of the team. We weren't impacted by the Atlassian outage, but Linear does seem to have a pretty compelling feature-set.

Re: Inside the longest Atlassian outage

#733

Earlier quoted context omitted.

They claim they test backups quarterly yet they don't have a procedure in place to restore the operation. We all know your backup is not tested until you restored everything successfully. This is not an engineering mistake, it is a flat out lie.

Well, their explanation makes sense. These are multi-tenant environments where not every tenant was affected; sensibly, the backups appear divided by environment, not tenant. You can’t blindly revert to an environment’s last backup in this scenario, although you’d think they would have done it before.

Of course you can't blindly restore, but it seems that's what they 'test'. Either they are completely incompetent or they don't test the real procedure.

Re: Inside the longest Atlassian outage

#734

Honest question here: The companies impacted by this, are they not taking backups of their Jira/Confluence/Bitbucket instances? Or is this outage impacting the ability to import those backups? There are some Python scripts that will back up Jira and Confluence. I whipped up a quick script that gets a list of all our bitbucket repos and then it clones those daily as well.

A Jira backup blob isn't especially useful. Confluence could be if it's essentially HTML dumps that you could host internally read-only. Bitbucket clearly has backups and a migration path.

Why is it not useful? From eyeballing it it looked like the same file I built from our Fogbugz data to import our historic cases into Jira. I'll carve out some time to try doing an import into a new project to see if it loads properly.

Re: Inside the longest Atlassian outage

#735

Engineering mistakes happen. The most inexcusable thing is not communicating with the paying customers who have been affected for over a week. Atlassian's Global Head of Customer Success probably should have been fired but here she is promoting Atlassian Cloud on LinkedIn three days ago: https://www.linkedin.com/mwlite/in/gertie-rizzo-5b70061 Actually reading a bit more, it seems like their customer team was partying…

A lot of the replies to this thread seem to be running with their own assumptions & definitions of what a "Customer Success" dept/org is. I've personally never worked for a company that used that language, so I went looking and found the job description for this person's role: https://startup.jobs/global-head-of-customer-success-remote-...

I'm still unclear where "Global Head of" for something like this fits in an org chart (Who do they report to? Who reports to them? Is it within a marketing, sales or customer support tree? Etc). Title inflation and whatnot considered...

Re: Inside the longest Atlassian outage

#736

I've repeatedly asked Atlassian if: 1. They can confirm that they have backups of our data (about a thousand stories, substantial confluence, opsgenie history, and three service desks). 2. Will our integrations, configuration, and customizations also be recovered, or will we need to rebuild those once our data is recovered? I have received no response, and no human is even willing to acknowledge those questions. The…

There is a thing that I dont understand, from their blog/report [1]

If the script was used in "permanently delete" mode, which is intended for compliance... how do you restore?

Is it the only explanation... if the deletion is non-compliant?

> Second, the script we used provided both the "mark for deletion" capability used in normal day-to-day operations (where recoverability is desirable), and the "permanently delete" capability that is required to permanently remove data when required for compliance reasons.

[1] https://www.atlassian.com/engineering/april-2022-outage-upda...

Re: Inside the longest Atlassian outage

#737

I've repeatedly asked Atlassian if: 1. They can confirm that they have backups of our data (about a thousand stories, substantial confluence, opsgenie history, and three service desks). 2. Will our integrations, configuration, and customizations also be recovered, or will we need to rebuild those once our data is recovered? I have received no response, and no human is even willing to acknowledge those questions. The…

I was down, my instance is fully restored right now. 1. They do, every 25 hours, via snapshot. I have spoked to their team since the incident and that same thing is in this article. 2. Yes, they recover all of it. Some things have had issues, external mailboxes attached to service management projects, some attachment rendering slowness. Filters needing to be overlayed into our instance again, but otherwise it is runn…

Thank you for responding here. Even a single successful restoration, anecdotal as it is, makes me feel a lot better about the situation. I was honestly wondering if Atlassian was just pushing the date out to soften the eventual backlash when they had to announce data loss.

Re: Inside the longest Atlassian outage

#738

What's a good Jira replacement? Redmine? Phabricator? OpenProject? Just leaving the jira server alone and hoping there's no new and exciting zero-days? One thing is clear, these guys are a bunch of cowboys who can't be trusted with any amount of data.

I loved Phabricator but it's end of life. https://www.phacility.com/phabricator/

Re: Inside the longest Atlassian outage

#739

I remember finding out one of the senior managers from my company ended up as head of software at Atlassian. It was at that point I was convinced Atlassian has no idea what the hell they're doing. I think this demonstrates the point nicely.

Did you fire them or they quit to take the opportunity? Unless it's the former you're just as guilty :)

Re: Inside the longest Atlassian outage

#740

> it takes between 4 and 5 elapsed days to hand a site back to a customer. Atlassian's SLA page says, Premium Cloud Products 99.9% That's 43 minutes of downtime per month. That works out to, Atlassian can't have any more downtime for the next 14 years. Are SLAs even real? I'm being slightly facetious. From the page text it's just a threshold after which I think you're entitled to some money back for that month.

"Shit happens" is a universal when it comes to computing. SLAs describe what is a normal background level of shit happening vs. what demands immediate attention and action from the team.
Post reply on HN