Live data from Hacker News

Inside the longest Atlassian outage

newsletter.pragmaticengineer.com

481–490 of 772 posts

Re: Inside the longest Atlassian outage

#482

Engineering mistakes happen. The most inexcusable thing is not communicating with the paying customers who have been affected for over a week. Atlassian's Global Head of Customer Success probably should have been fired but here she is promoting Atlassian Cloud on LinkedIn three days ago: https://www.linkedin.com/mwlite/in/gertie-rizzo-5b70061 Actually reading a bit more, it seems like their customer team was partying…

No post body was provided.

Re: Inside the longest Atlassian outage

#483
post #2

Selectively restoring data only for certain rows is super hard. But the communications by Atlassian has been the worst I have ever seen in the industry.

Hi, this is Mike from Atlassian Engineering. You are right the communications from us have not lived up to our standard. We will focus on this specifically once we restore service and get the post incident review out there. More details here: https://www.atlassian.com/engineering/april-2022-outage-upda...

Spamming HN isn't helping your cause man.

Re: Inside the longest Atlassian outage

#484
post #442
post #411

This is extremely poor for a large SaaS company. A standard RFP question for SaaS should be: - Can you restore data for a single customer, and if so, what is the RTO for that operation? A smaller SaaS could be excused for only thinking about full database restores. When you're a scrappy upstart, thinking about hypotheticals is less important than survival. But for any decent size multi-tenanted SaaS, it's imperative…

Which SaaS platforms provide account-level restores? If you contact them and say "please restore our data to as it was last week" those I know do not offer this.

I accidentally built out this feature at a company once and it totally saved our asses a week later.

Re: Inside the longest Atlassian outage

#485
post #367

Earlier quoted context omitted.

Poor taste, buddy. Comparing the Atlassian mess-up to the Holocaust diminishes the Holocaust.

um... the sentiment is universal it's not specific to that particularly awful history. Sorry if it triggered you, HN doesn't offer a delete button. FYI my ancestors fled oppression on both sides and I'm well aware that it's a miracle I'm alive. Again, one bad thing leading to another is a common human behavior, and the Holocaust is just an extreme example that I ABSOLUTELY did not intend whatsoever. You make this con…

If I'm then one making this connection, then it should be trivial for you to finish your sentence. "First they came for..." Who are the Jews in your analogy ? Who are the communists? the trade unionists? And who is the totalitarian regime?

Suggesting that Niemöller's poem is about "one bad thing leads to another" is like suggesting that Anne Frank's diary is about "sometimes girls have really bad days." I understand you didn't mean any offense to anyone. But that's not a license to be offensive, and then duck for cover.

Re: Inside the longest Atlassian outage

#486
post #480

Engineering mistakes happen. The most inexcusable thing is not communicating with the paying customers who have been affected for over a week. Atlassian's Global Head of Customer Success probably should have been fired but here she is promoting Atlassian Cloud on LinkedIn three days ago: https://www.linkedin.com/mwlite/in/gertie-rizzo-5b70061 Actually reading a bit more, it seems like their customer team was partying…

sales never takes the blame. If anyone is fired it will be scapegoats in engineering once they have busted their ass to restore their reward will be the door

This is an engineering problem. They should own it and improve things, make sure it doesn't happen again.

Also, GP's quote

> Engineering mistakes happen.

I don't like this statement because it offers consolation at the expense of unintentional normalization.

Re: Inside the longest Atlassian outage

#487

Earlier quoted context omitted.

What's the difference?

Good faith would be to lose all of that money to people who are already your customers. Business-wise would be to stay in their good graces and keep those customers by offering the refund, but you don't lose any money to those who either don't care or won't move to a competitor.

25 years ago the clutch in my beater truck was slipping. I was 16 years old, making $50 a week and had very little in savings. I took that truck to a shop within walking distance of my job.

2 hours later I walked back to see what they found. I figured it would be several hundred dollars for a new clutch, and I'd have to borrow money or something to get it done. I talked to the owner who told be it was an adjustment on the cable. Just needed to be scootched up a bit and it was probably good for another 30k miles.

When I asked him how much I owed, he laughed at me and said, "For that? Not worth writing it up. No charge. You want me to show you how to do it yourself next time?"

The shop could very easily have charged me 1 hour of labor at their standard rate, maybe $75 or so. Plus a diagnostic or test drive fee. Whatever. He could have told me, "$123.98" and I would have paid it. I wouldn't even have been mad. But I sure as hell wouldn't have remembered the experience so clearly. Nor would I have told a dozen people over the years to take their cars there. And I definitely would not have driven 20 miles out of my way to return to that shop in the future years.

Being cynical about this stuff will hurt your brand. It's not obvious. It doesn't show up on the earnings report as a line item. This is service segmentation that seems like a no-brainer to a clueless MBA, but actually matters in the long run. How people view your brand is immensely important.

Not forcing customers you already screwed over to then spend more time chasing a refund is not only the right thing to do, it's also good business.

Re: Inside the longest Atlassian outage

#488
post #385
post #142

We use on-premises setups for almost everything (we generally avoid cloud solutions to have full control of our data), sometimes (approximately once a month) it goes down for a few minutes which already feels like a torture because all our processes depend on it, I can't imagine having no access to it for several weeks, all our work would stop to a halt... The office of the guy who administers on-premise servers is l…

What do you do if your on prem setup lost data? There is an implicit assumption here that on prem is more reliable than cloud. Less downtime, less chances of data loss etc. Obviously it depends on which cloud product we're talking about but I don't think a blanket "my on prem goes down less and when it does go down I can get it back up sooner" is true.

I think for that question we also have to define on-prem just to be clear. To many on-prem means "own cloud subscription".

Re: Inside the longest Atlassian outage

#489
post #480

Earlier quoted context omitted.

sales never takes the blame. If anyone is fired it will be scapegoats in engineering once they have busted their ass to restore their reward will be the door

This is an engineering problem. They should own it and improve things, make sure it doesn't happen again. Also, GP's quote > Engineering mistakes happen. I don't like this statement because it offers consolation at the expense of unintentional normalization.

And coders that say all code has bugs are just defeatists that are trying to make excuses for being lazy.

Sometimes manure will always hit the fan. Being robust means being able to handle that.

Re: Inside the longest Atlassian outage

#490

What's a good Jira replacement? Redmine? Phabricator? OpenProject? Just leaving the jira server alone and hoping there's no new and exciting zero-days? One thing is clear, these guys are a bunch of cowboys who can't be trusted with any amount of data.

I've used Request Tracker for years. It's not pretty, it's written in Perl, but I can fairly easily make it do all the ticket tracking flows I care about and it just runs and runs and runs. My scale is admittedly small, but I put tens of thousands of tickets per year through my instance, and i basically never have to touch it unless I'm setting up a new queue or different flow for something.

Wow, I’ve never seen anyone mention RT here. I used it for years when I was working IT for my university while in undergrad. It worked pretty well. It didn’t have a lot of features but it allowed clients/customers to respond to tickets via email which was pretty cool at the time (late 00s). It also ran pretty fast on the terrible servers we had it on.
Post reply on HN