Live data from Hacker News

Inside the longest Atlassian outage

newsletter.pragmaticengineer.com

761–770 of 772 posts

Re: Inside the longest Atlassian outage

#761

Earlier quoted context omitted.

There is irony in complaining about over-communication when it's in response to criticisms of under-communication.

Key word "spamming." It wasn't communication but another dry and information-free blob of text. Communication requires something to say.

It's worse than that, they're saying communication was not up to their standards without actually communicating anything we didn't already know.

At least explain why there was such a total communication blackout company wide. Even support staff weren't allowed to discus it. Why?

Re: Inside the longest Atlassian outage

#762

Earlier quoted context omitted.

We migrated from Slack to self-hosted Mattermost so we avoid being down. (And I guess money.) Mattermost is so much worse that the slowness and general issues are not worth it. And in the end it is more down than Slack ever was, because it has performance issues. I am not sure if it is Mattermost fault or our fault; but my friend from other corporation has similar experience with it. But maybe in general just don't k…

We still haven't taken IRC down because it's our backup for when slack goes down. I swear if IRC just implemented emojis.

we have WhatsApp as Mattermost backup.

Lately we use it more than mattermost :)

Re: Inside the longest Atlassian outage

#763

Earlier quoted context omitted.

Why is it not useful? From eyeballing it it looked like the same file I built from our Fogbugz data to import our historic cases into Jira. I'll carve out some time to try doing an import into a new project to see if it loads properly.

In this particular case, if you have another account, that could work. What I meant was that Jira isn't very useful unless people can actively use it for issue tracking. It's not all that valuable when it's just a reference.

Ooh, I see. In my case I'm thinking more that in the case of Atlassian having lost our data we could reload it from that backup.

Re: Inside the longest Atlassian outage

#764

Earlier quoted context omitted.

I've used Request Tracker for years. It's not pretty, it's written in Perl, but I can fairly easily make it do all the ticket tracking flows I care about and it just runs and runs and runs. My scale is admittedly small, but I put tens of thousands of tickets per year through my instance, and i basically never have to touch it unless I'm setting up a new queue or different flow for something.

Wow, I’ve never seen anyone mention RT here. I used it for years when I was working IT for my university while in undergrad. It worked pretty well. It didn’t have a lot of features but it allowed clients/customers to respond to tickets via email which was pretty cool at the time (late 00s). It also ran pretty fast on the terrible servers we had it on.

We still run it today; they had a major release last year, I think. Its key feature is that it remains email-first. Customers never interface with the website, for them it's all just like they're emailing a human, with some extra tooling and tracking on top.

Re: Inside the longest Atlassian outage

#765
post #245

Earlier quoted context omitted.

Is it? We use JIRA. Not impacted. If this had hit us.. we would just switch to excel or something for a week/month? But maybe we are a very light user of JIRA. Nothing in there can't be replaced. It's "nice" to be able to go look up a 3 year old bug and which client reported it, but not really crucial for day to day ops.

"Switch to Excel for a week/month" Right.

Excuse me? THis is my company I founded with my own 2 hands.

We used excel for the first year. Trello for the second.

I know what we need.

Re: Inside the longest Atlassian outage

#766
post #96

Earlier quoted context omitted.

Is it? We use JIRA. Not impacted. If this had hit us.. we would just switch to excel or something for a week/month? But maybe we are a very light user of JIRA. Nothing in there can't be replaced. It's "nice" to be able to go look up a 3 year old bug and which client reported it, but not really crucial for day to day ops.

I wonder why you use Jira if a spreadsheet is sufficient for your use case.

It allows some simple links like story X depends on story Y, and displays it in a visually pleasent way.

Re: Inside the longest Atlassian outage

#767

Earlier quoted context omitted.

Hi, I'm Mike and I work in Engineering at Atlassian. Here's our approach to backup and data management: https://www.atlassian.com/trust/security/data-management - we certainly have the backups and have a restore process that we keep to. However, this incident stressed our ability to do this at scale, which has led to the very long times to restore.

Hey Mike; Not dumping on you personally, but the RTO claims to be 6 hours. I can understand that being a target, but we're at 32X that RTO target, with a communicated target date of another 12 or so days IIRC. That's literally two orders of magnitude longer than the RTO. I don't think any rational person would take that document seriously at this point. I'll also ask (since nobody else has answered, I may as well ask…

Hi Ranteki, you're right that the RTO for this incident is far longer than any of the ones listed on the doc I linked above. That's because our RPO/RTO targets are set at the service level and not at the level of a "customer". This is part of the problem and demonstrates a gap both in what the doc is meant to express and a gap in our automation. Both will be reviewed in the PIR. Also, the answer to (1) and (2) is yes.

Re: Inside the longest Atlassian outage

#768

Regarding the backup restores: I once worked a company that had a data loss issue. There was nothing else we could do, we had exhausted every option we had over almost 40 hours. At the end of the second day, it was decided to restore from backup. We had done this before, as a test. It took about 12 hours to restore the data and another 12 hours to import the data and get back up and running. One small thing was diffe…

Did I ever tell you guys about the time we accidentally nuked all the mailboxes for all the million-plus users on The Global Network Navigator (GNN) site? And how the restore process failed for us?

This hasn't been written up at The Register yet, so I don't have a single URL I can share with you.

Re: Inside the longest Atlassian outage

#769

Earlier quoted context omitted.

Hi, I'm Mike and I work in Engineering at Atlassian. Here's our approach to backup and data management: https://www.atlassian.com/trust/security/data-management - we certainly have the backups and have a restore process that we keep to. However, this incident stressed our ability to do this at scale, which has led to the very long times to restore.

A friend in Atlassian engineering said the numbers on the trust site are closer to wishful thinking than actual capabilities, and that there has been an engineering wide disaster recovery project running because things were in such bad shape. The recovery part hasn't even started. If Atlassian could actually restore full products in under six hours, they should have been able to restore a second copy of the products…

Nah. The RTO/RPO assumes that only one customer that has a failure big enough to require a restore.

When the entire service is hosed, that's a totally different set of circumstances, and you have to look at what the RTO/RPO are for basically restoring the entire service for all customers. And since the have more than a thousand customers, it totally makes sense that it would take orders of magnitude longer to restore the entire service.

Re: Inside the longest Atlassian outage

#770
post #266

All I can say as an Attlassian Server products user is that the moment they say it was Cloud or nothing, I choose nothing. I much rather running Gittea on a raspberry pi that I CONTROL than having to have the impotence of doing nothing for more than a week. + having work at cloud companies and having been requested to "collect customer data" to hand it over to the government I would NEVER move critical pieces to anyo…

gittea is nice, but for more advanced project management with a self-hosted option, check out youtrack
Post reply on HN