Earlier quoted context omitted.
Nope. I exported our data after they restored the backup and then we cancelled less than a month later. Like I obviously understand suspending our logins, but why would you ever delete someone's data when it's literally only 160 KB of text? The whole thing made zero sense.
> why would you ever delete someone's data when it's literally only 160 KB of text? Compliance? The contract has expired, so there’s no legal basis for them to keep your data?
Inside the longest Atlassian outage
691–700 of 772 posts
Re: Inside the longest Atlassian outage
#692Earlier quoted context omitted.
People who call for other people's firings in organizations that they have no visibility into are so weird. This post reads like a tech outage's version of cancel culture where trying to find someone to blame and skewer for an injustice is more important than actually determining how much (if any) blame they deserve for it Also posting a LinkedIn event photos with people's real names and pictures in a top post on HN…
I agree. Customer success is a support role, this was an engineering mistake. Can't blame support for something an engineer did.
Re: Inside the longest Atlassian outage
#693This is extremely poor for a large SaaS company. A standard RFP question for SaaS should be: - Can you restore data for a single customer, and if so, what is the RTO for that operation? A smaller SaaS could be excused for only thinking about full database restores. When you're a scrappy upstart, thinking about hypotheticals is less important than survival. But for any decent size multi-tenanted SaaS, it's imperative…
I still don't get why they didn't separate clients on a database level. Sure, put many clients on one database server to save resources. But why not use different databases? They cost nothing and provide perfect separation. It also drastically lowers the attack surface as you can set all permissions via database software. And if they had done that, this would've never been a multi-day outage. If Jira was a product us…
Really annoying things that slash your velocity. You can't easily run pan-customer queries, can't aggregate data.
Ironically too using a single database makes full backup and restore much easier.
Re: Inside the longest Atlassian outage
#694Earlier quoted context omitted.
I am tired of survivor-biased "best practices" advice. I wonder which practices contained there are the worst practices .
This is interesting but can you expand? My understanding of survivor bias is that you're getting a skewed picture because some of the data was excluded completely.
It is only through understanding what can fail that you can figure out causation.
And since Atlassian failed here, the talk might expose some of the failure's causes, or at least cast doubt over the usefulness of the practices presented.
Re: Inside the longest Atlassian outage
#695Asking for a friend.
Re: Inside the longest Atlassian outage
#696This is extremely poor for a large SaaS company. A standard RFP question for SaaS should be: - Can you restore data for a single customer, and if so, what is the RTO for that operation? A smaller SaaS could be excused for only thinking about full database restores. When you're a scrappy upstart, thinking about hypotheticals is less important than survival. But for any decent size multi-tenanted SaaS, it's imperative…
I still don't get why they didn't separate clients on a database level. Sure, put many clients on one database server to save resources. But why not use different databases? They cost nothing and provide perfect separation. It also drastically lowers the attack surface as you can set all permissions via database software. And if they had done that, this would've never been a multi-day outage. If Jira was a product us…
I understand the sentiment, but This is a pretty simplistic take that I very much doubt will hold true for meaningful traffic. Many databases have licensing considerations that arent amenable. Beyond that you get in to density and resource problems as simple as IO, processes, threads etc. But most of all theres the time and effort burden in supporting migrations, schema updates, etc.
Yes layered logical separation is a really good idea. Its also really expensive once you start dealing with organic growth and a meaningful number of discrete customers.
Disclaimer: Principal at AWS who was helped build and run services with both multi tenant and single tenant architectures.
Re: Inside the longest Atlassian outage
#697Earlier quoted context omitted.
I still don't get why they didn't separate clients on a database level. Sure, put many clients on one database server to save resources. But why not use different databases? They cost nothing and provide perfect separation. It also drastically lowers the attack surface as you can set all permissions via database software. And if they had done that, this would've never been a multi-day outage. If Jira was a product us…
> why not use different databases? They cost nothing and provide perfect separation. I understand the sentiment, but This is a pretty simplistic take that I very much doubt will hold true for meaningful traffic. Many databases have licensing considerations that arent amenable. Beyond that you get in to density and resource problems as simple as IO, processes, threads etc. But most of all theres the time and effort bu…
And for migrations and schema updates I'd see this as a huge advantage. Migrating customers one by one is much easier than everyone at once. You also never have the issue that operations at one customer could cause a global lock affecting other customers.
Of course resource sharing isn't easy in this scenario, but you'd never want to connect data between customers anyway so I don't see the issue with that.
But maybe it works harder in a cloud environment where more is abstracted away.
Re: Inside the longest Atlassian outage
#698Engineering mistakes happen. The most inexcusable thing is not communicating with the paying customers who have been affected for over a week. Atlassian's Global Head of Customer Success probably should have been fired but here she is promoting Atlassian Cloud on LinkedIn three days ago: https://www.linkedin.com/mwlite/in/gertie-rizzo-5b70061 Actually reading a bit more, it seems like their customer team was partying…
https://www.linkedin.com/feed/update/urn:li:activity:6918235...
Re: Inside the longest Atlassian outage
#699Engineering mistakes happen. The most inexcusable thing is not communicating with the paying customers who have been affected for over a week. Atlassian's Global Head of Customer Success probably should have been fired but here she is promoting Atlassian Cloud on LinkedIn three days ago: https://www.linkedin.com/mwlite/in/gertie-rizzo-5b70061 Actually reading a bit more, it seems like their customer team was partying…
These are base instincts speaking. Emotional, vengeful, a desire to punish because you've been hurt. It's perfectly natural to feel that way, but it's not reasonable, fair, or effective.
However it doesn’t imply there is vindictive drive
some people will not have a good month career wise, some people will lose trust, some people will be an consequences of regaining the public confidence
Re: Inside the longest Atlassian outage
#700Engineering mistakes happen. The most inexcusable thing is not communicating with the paying customers who have been affected for over a week. Atlassian's Global Head of Customer Success probably should have been fired but here she is promoting Atlassian Cloud on LinkedIn three days ago: https://www.linkedin.com/mwlite/in/gertie-rizzo-5b70061 Actually reading a bit more, it seems like their customer team was partying…
People who call for other people's firings in organizations that they have no visibility into are so weird. This post reads like a tech outage's version of cancel culture where trying to find someone to blame and skewer for an injustice is more important than actually determining how much (if any) blame they deserve for it Also posting a LinkedIn event photos with people's real names and pictures in a top post on HN…
That’s my perspective. Reading your comment it feels like between the lines you don’t think this is as serious as other people. If you have a different perspective, could you come out and say it? Can you elaborate on why a “Global Head Of” isn’t a senior leader? Do you think this is an unfortunate tech outage that is to be expected from any b2b tech company of this size? If you did, I wonder how many people here would disagree with you. Implicitly, the person to whom you are responding does. “Business collapse” and “vegas party” do not look good getting caught in bed together.
Your point would come across much better if it wasn’t mixed with moral outrage. If you have alternative opinions please share them and back them up. Then you will have earned a little more of the massive amount of social capital you need to tell someone, in public and quite rudely, to do better.