Live data from Hacker News

Inside the longest Atlassian outage

newsletter.pragmaticengineer.com

491–500 of 772 posts

Re: Inside the longest Atlassian outage

#491

Earlier quoted context omitted.

With GDPR, privacy regulations and data breach regulations sweeping the globe, holding onto unnecessary data is a huge liability. Getting rid of data you no longer have clear consent to store, or which you're unlikely to have a clear business need to continue storing, is a sign of a good company these days.

True, but likely not this kind of data.

Yes, this kind of data. Your OkCupid account has all kinds of information about who you associate with.

Re: Inside the longest Atlassian outage

#492

Engineering mistakes happen. The most inexcusable thing is not communicating with the paying customers who have been affected for over a week. Atlassian's Global Head of Customer Success probably should have been fired but here she is promoting Atlassian Cloud on LinkedIn three days ago: https://www.linkedin.com/mwlite/in/gertie-rizzo-5b70061 Actually reading a bit more, it seems like their customer team was partying…

Can confirm. Saw them there while I was on vacation.

Jesus, if there was ever an example of the internet making the world smaller.

When do execs living it up at the fucking Wynn Encore while the house burns down start to not get another job?

They’ll keep pulling this shit until it cost money.

Re: Inside the longest Atlassian outage

#493

Earlier quoted context omitted.

This is an engineering problem. They should own it and improve things, make sure it doesn't happen again. Also, GP's quote > Engineering mistakes happen. I don't like this statement because it offers consolation at the expense of unintentional normalization.

And coders that say all code has bugs are just defeatists that are trying to make excuses for being lazy. Sometimes manure will always hit the fan. Being robust means being able to handle that.

A culture where mistakes are taken too seriously or too lightly leads to problems. Also it depends on what stage of the product cycle (Innovation/Rapid Development vs. Robustness/Quality). I'd argue that Atlassian products should err towards robustness and high quality. Not trying to break any new ground.

Re: Inside the longest Atlassian outage

#494

Earlier quoted context omitted.

With GDPR, privacy regulations and data breach regulations sweeping the globe, holding onto unnecessary data is a huge liability. Getting rid of data you no longer have clear consent to store, or which you're unlikely to have a clear business need to continue storing, is a sign of a good company these days.

True, but likely not this kind of data.

No post body was provided.

Re: Inside the longest Atlassian outage

#495

so this is the end of Atlassian as a company right?

I had the same initial thought. Surely a weekslong outage would drive customers away permanently, right? Nope. From TFA: > I asked customers if they would offboard Atlassian as a result of the outage. Most of them said they won’t leave the Atlassian stack, as long as they don’t lose data. This is because moving is complex and they don’t see a move would mitigate a risk of a cloud provider going down.

it doesn't happen overnight, but this is a really bad precedent and it will definitely have an effect on both sales AND renewals. This market is theirs to lose and seems they are doing everything they can to do just that. Github is getting better, and it has mindshare amongst developers, not to mention it's part of a company that like it or not knows how to sell to large enterprises (Microsoft).

Re: Inside the longest Atlassian outage

#496

Earlier quoted context omitted.

Can confirm. Saw them there while I was on vacation.

Jesus, if there was ever an example of the internet making the world smaller. When do execs living it up at the fucking Wynn Encore while the house burns down start to not get another job? They’ll keep pulling this shit until it cost money.

For clarity: I went through a period where some combination of self-indulgence and legitimate life crisis caused me to take my eye off the ball when it mattered.

I’m still trying to kickstart a second act years later, because I’m trailer trash and it’s hard work when you’re that.

Re: Inside the longest Atlassian outage

#497

> it takes between 4 and 5 elapsed days to hand a site back to a customer. Atlassian's SLA page says, Premium Cloud Products 99.9% That's 43 minutes of downtime per month. That works out to, Atlassian can't have any more downtime for the next 14 years. Are SLAs even real? I'm being slightly facetious. From the page text it's just a threshold after which I think you're entitled to some money back for that month.

Hi, this is Mike from Atlassian Engineering. For the customers impacted by this incident covered by an SLA, we will adhere to our contractual terms. However, given the long duration of this outage, we are planning to go above and beyond for our impacted customers. We are currently focused on restoring service, but after that will be discussing how we can make it right for each impacted customer.

It looks like you are focused on Hacker News comments.

Re: Inside the longest Atlassian outage

#498

> Most of them said they won’t leave the Atlassian stack, as long as they don’t lose data. This is because moving is complex and they don’t see a move would mitigate a risk of a cloud provider going down. I still don't understand the strangehold JIRA has on some clients. I can't quickly think of another SaaS product that could be down for almost 2 weeks and not have most customers leave.

Atlassian sells to execs and gives kickbacks. You don't want to burn the company that gave you money and that you pushed through although you knew they sucked.

Re: Inside the longest Atlassian outage

#499

Gmail had a vaguely similar outage years ago. [1] tl;dr: 1. Different root cause. There was a bug in a refactoring of gmail's storage layer (iirc a missing asterisk caused a pointer to an important bool to be set to null, rather than setting the bool to false), which slipped through code review, automated testing, and early test servers dedicated to the team, so it got rolled out to some fraction of real users. Onlin…

Funny enough, most of what we restored then was spam (ex gTape SRE, remember the outage).

Re: Inside the longest Atlassian outage

#500
post #480

Earlier quoted context omitted.

sales never takes the blame. If anyone is fired it will be scapegoats in engineering once they have busted their ass to restore their reward will be the door

This is an engineering problem. They should own it and improve things, make sure it doesn't happen again. Also, GP's quote > Engineering mistakes happen. I don't like this statement because it offers consolation at the expense of unintentional normalization.

Ever heard of Space Shuttle Challenger? You cant own it if your management is against it.
Post reply on HN