Inside the longest Atlassian outage
481–490 of 772 posts
Re: Inside the longest Atlassian outage
#482Engineering mistakes happen. The most inexcusable thing is not communicating with the paying customers who have been affected for over a week. Atlassian's Global Head of Customer Success probably should have been fired but here she is promoting Atlassian Cloud on LinkedIn three days ago: https://www.linkedin.com/mwlite/in/gertie-rizzo-5b70061 Actually reading a bit more, it seems like their customer team was partying…
Re: Inside the longest Atlassian outage
#483Selectively restoring data only for certain rows is super hard. But the communications by Atlassian has been the worst I have ever seen in the industry.
Hi, this is Mike from Atlassian Engineering. You are right the communications from us have not lived up to our standard. We will focus on this specifically once we restore service and get the post incident review out there. More details here: https://www.atlassian.com/engineering/april-2022-outage-upda...
Re: Inside the longest Atlassian outage
#484This is extremely poor for a large SaaS company. A standard RFP question for SaaS should be: - Can you restore data for a single customer, and if so, what is the RTO for that operation? A smaller SaaS could be excused for only thinking about full database restores. When you're a scrappy upstart, thinking about hypotheticals is less important than survival. But for any decent size multi-tenanted SaaS, it's imperative…
Which SaaS platforms provide account-level restores? If you contact them and say "please restore our data to as it was last week" those I know do not offer this.
Re: Inside the longest Atlassian outage
#485Earlier quoted context omitted.
Poor taste, buddy. Comparing the Atlassian mess-up to the Holocaust diminishes the Holocaust.
um... the sentiment is universal it's not specific to that particularly awful history. Sorry if it triggered you, HN doesn't offer a delete button. FYI my ancestors fled oppression on both sides and I'm well aware that it's a miracle I'm alive. Again, one bad thing leading to another is a common human behavior, and the Holocaust is just an extreme example that I ABSOLUTELY did not intend whatsoever. You make this con…
Suggesting that Niemöller's poem is about "one bad thing leads to another" is like suggesting that Anne Frank's diary is about "sometimes girls have really bad days." I understand you didn't mean any offense to anyone. But that's not a license to be offensive, and then duck for cover.
Re: Inside the longest Atlassian outage
#486Engineering mistakes happen. The most inexcusable thing is not communicating with the paying customers who have been affected for over a week. Atlassian's Global Head of Customer Success probably should have been fired but here she is promoting Atlassian Cloud on LinkedIn three days ago: https://www.linkedin.com/mwlite/in/gertie-rizzo-5b70061 Actually reading a bit more, it seems like their customer team was partying…
sales never takes the blame. If anyone is fired it will be scapegoats in engineering once they have busted their ass to restore their reward will be the door
Also, GP's quote
> Engineering mistakes happen.
I don't like this statement because it offers consolation at the expense of unintentional normalization.
Re: Inside the longest Atlassian outage
#487Earlier quoted context omitted.
What's the difference?
Good faith would be to lose all of that money to people who are already your customers. Business-wise would be to stay in their good graces and keep those customers by offering the refund, but you don't lose any money to those who either don't care or won't move to a competitor.
2 hours later I walked back to see what they found. I figured it would be several hundred dollars for a new clutch, and I'd have to borrow money or something to get it done. I talked to the owner who told be it was an adjustment on the cable. Just needed to be scootched up a bit and it was probably good for another 30k miles.
When I asked him how much I owed, he laughed at me and said, "For that? Not worth writing it up. No charge. You want me to show you how to do it yourself next time?"
The shop could very easily have charged me 1 hour of labor at their standard rate, maybe $75 or so. Plus a diagnostic or test drive fee. Whatever. He could have told me, "$123.98" and I would have paid it. I wouldn't even have been mad. But I sure as hell wouldn't have remembered the experience so clearly. Nor would I have told a dozen people over the years to take their cars there. And I definitely would not have driven 20 miles out of my way to return to that shop in the future years.
Being cynical about this stuff will hurt your brand. It's not obvious. It doesn't show up on the earnings report as a line item. This is service segmentation that seems like a no-brainer to a clueless MBA, but actually matters in the long run. How people view your brand is immensely important.
Not forcing customers you already screwed over to then spend more time chasing a refund is not only the right thing to do, it's also good business.
Re: Inside the longest Atlassian outage
#488We use on-premises setups for almost everything (we generally avoid cloud solutions to have full control of our data), sometimes (approximately once a month) it goes down for a few minutes which already feels like a torture because all our processes depend on it, I can't imagine having no access to it for several weeks, all our work would stop to a halt... The office of the guy who administers on-premise servers is l…
What do you do if your on prem setup lost data? There is an implicit assumption here that on prem is more reliable than cloud. Less downtime, less chances of data loss etc. Obviously it depends on which cloud product we're talking about but I don't think a blanket "my on prem goes down less and when it does go down I can get it back up sooner" is true.
Re: Inside the longest Atlassian outage
#489Earlier quoted context omitted.
sales never takes the blame. If anyone is fired it will be scapegoats in engineering once they have busted their ass to restore their reward will be the door
This is an engineering problem. They should own it and improve things, make sure it doesn't happen again. Also, GP's quote > Engineering mistakes happen. I don't like this statement because it offers consolation at the expense of unintentional normalization.
Sometimes manure will always hit the fan. Being robust means being able to handle that.
Re: Inside the longest Atlassian outage
#490What's a good Jira replacement? Redmine? Phabricator? OpenProject? Just leaving the jira server alone and hoping there's no new and exciting zero-days? One thing is clear, these guys are a bunch of cowboys who can't be trusted with any amount of data.
I've used Request Tracker for years. It's not pretty, it's written in Perl, but I can fairly easily make it do all the ticket tracking flows I care about and it just runs and runs and runs. My scale is admittedly small, but I put tens of thousands of tickets per year through my instance, and i basically never have to touch it unless I'm setting up a new queue or different flow for something.