Live data from Hacker News

Inside the longest Atlassian outage

newsletter.pragmaticengineer.com

751–760 of 772 posts

Re: Inside the longest Atlassian outage

#751

Earlier quoted context omitted.

Not every business can afford to go one month without income. What's the best thing for customers? Have the business go bankrupt and irremediably lose access to the service?

It's 400 clients, not all their user base. They can handle the lost income from a small slice of their customers for one month. And if they can't sustain that, then it's even more imperative that those customers migrate away.

Fastmail gave 1 month free service to about 2/3 of our customers after a major disk failure that led to about a week of downtime for them as we recovered from backups in... 2005ish I think. Long time ago - it was a pretty major hit and the wave in income is still visible all these years later as a lean month where there's no renewals from that batch! Definitely the right thing to do though.

Re: Inside the longest Atlassian outage

#752

Earlier quoted context omitted.

A Jira backup blob isn't especially useful. Confluence could be if it's essentially HTML dumps that you could host internally read-only. Bitbucket clearly has backups and a migration path.

Why is it not useful? From eyeballing it it looked like the same file I built from our Fogbugz data to import our historic cases into Jira. I'll carve out some time to try doing an import into a new project to see if it loads properly.

In this particular case, if you have another account, that could work. What I meant was that Jira isn't very useful unless people can actively use it for issue tracking. It's not all that valuable when it's just a reference.

Re: Inside the longest Atlassian outage

#753

Earlier quoted context omitted.

I agree. Customer success is a support role, this was an engineering mistake. Can't blame support for something an engineer did.

Seems like you're still assigning blame. Incidents are rarely if ever monocausal. The fantastic and accurate point the GP made is that fingerpointing is pointless. Much better to seek to learn and understand, which is always difficult but definitely can't be done from the sidelines without speaking to those involved.

Which is why I used mistake, the context is blaming support for an engineering action. Support is purely reactive.

Re: Inside the longest Atlassian outage

#754

Regarding the backup restores: I once worked a company that had a data loss issue. There was nothing else we could do, we had exhausted every option we had over almost 40 hours. At the end of the second day, it was decided to restore from backup. We had done this before, as a test. It took about 12 hours to restore the data and another 12 hours to import the data and get back up and running. One small thing was diffe…

This is why I love GCP Cloud Storage. The "colder" tiers are cheaper, and reads simply cost a lot more from there, but they don't slow them down and take days to restore. You pay with dollars, not time for restoring those GCS backups. e.g. Coldline [1] simply has reduced availability in exchange for being cheaper (99.9-99.95% availability, so 43min/mo, way less than "two days"). [1] https://cloud.google.com/storage/d…

Yea, till they rug pull it.

Re: Inside the longest Atlassian outage

#755

Engineering mistakes happen. The most inexcusable thing is not communicating with the paying customers who have been affected for over a week. Atlassian's Global Head of Customer Success probably should have been fired but here she is promoting Atlassian Cloud on LinkedIn three days ago: https://www.linkedin.com/mwlite/in/gertie-rizzo-5b70061 Actually reading a bit more, it seems like their customer team was partying…

When you depend on the clown, it will hit you some day. So if you are hit by such, now finally move away again from only depending on the clown, and (for example) re-value your on-prem-ity.

Re: Inside the longest Atlassian outage

#757
post #142

We use on-premises setups for almost everything (we generally avoid cloud solutions to have full control of our data), sometimes (approximately once a month) it goes down for a few minutes which already feels like a torture because all our processes depend on it, I can't imagine having no access to it for several weeks, all our work would stop to a halt... The office of the guy who administers on-premise servers is l…

There are always technical and economical pluses and minuses to any approach - but never underestimate the politics and character flaws that see dubious decisions pushed through by senior management regardless of those rational arguments pondered by the lower orders.

Re: Inside the longest Atlassian outage

#758

Do anyone seriously consider changing Jira/confluence to some alternative after this? I personally stopped using Jira a couple of years ago in projects I lead.

Some PMs at my company have been agitating for a switch from Asana to Jira, and now I (as VP Engineering) can guarantee that will never, ever happen.

What are PMs' arguments for the switch?

Re: Inside the longest Atlassian outage

#759
post #142

We use on-premises setups for almost everything (we generally avoid cloud solutions to have full control of our data), sometimes (approximately once a month) it goes down for a few minutes which already feels like a torture because all our processes depend on it, I can't imagine having no access to it for several weeks, all our work would stop to a halt... The office of the guy who administers on-premise servers is l…

We migrated from Slack to self-hosted Mattermost so we avoid being down. (And I guess money.) Mattermost is so much worse that the slowness and general issues are not worth it. And in the end it is more down than Slack ever was, because it has performance issues. I am not sure if it is Mattermost fault or our fault; but my friend from other corporation has similar experience with it. But maybe in general just don't k…

We still haven't taken IRC down because it's our backup for when slack goes down.

I swear if IRC just implemented emojis.

Re: Inside the longest Atlassian outage

#760

Earlier quoted context omitted.

As I understood it is not "Cloud or Nothing" but "Cloud or Data Center" - is this wrong?

We're a team with effectively is Cloud or nothing, and we weren't very keen on going Cloud even before this current clusterf...

Checkout the new GitHub projects boards / tables. More power than I need but it's closing the gap on some key features JIRA has.
Post reply on HN