Live data from Hacker News

Inside the longest Atlassian outage

newsletter.pragmaticengineer.com

771–772 of 772 posts

Re: Inside the longest Atlassian outage

#771
post #434

Earlier quoted context omitted.

We have all sorts of slack channels set up to coordinate activity, so that internal customers can talk to engineers easily, or engineers can engage with each other. If slack goes down, we'd have to work all that out. For many days, it would be a huge drag on the process, slowing down interactions. Other IM platforms wouldn't solve that just by existing. Sure, in principle one could set up such channels elsewhere, but…

Sounds like having a fallback pre-defined would be prudent if it's that important and you don't feel you could collectively extemporise something. "If Slack goes down, the plan is to use WhatsApp/Teams/Jeff's Matrix homeserver in his garage until service comes back. A list of group channels will be emailed if that happens." Then if it does go down, you don't have to waste the first day arguing about the plan.

People would forget, but having a well-documented plan at a well-understood URL would not be a bad idea.

Re: Inside the longest Atlassian outage

#772

I've repeatedly asked Atlassian if: 1. They can confirm that they have backups of our data (about a thousand stories, substantial confluence, opsgenie history, and three service desks). 2. Will our integrations, configuration, and customizations also be recovered, or will we need to rebuild those once our data is recovered? I have received no response, and no human is even willing to acknowledge those questions. The…

Update: My instance was recovered over this weekend (Apr 16). We've verified that integrations (webhooks, jira integrations) are working, and no data was lost that we can see (nobody works overnight, so our last change was at the end of the working day, which matched our recollection).
Post reply on HN