Live data from Hacker News

Atlassian: We estimate the rebuilding effort to last for up to 2 more weeks

twitter.com

211–220 of 260 posts

Re: Atlassian: We estimate the rebuilding effort to last for up to 2 more weeks

#211
post #78

Engineers, do yourself a favor and add to your CV/resume: I will not work for companies that use Atlassian software.

Yea uhh this kind of blanket statement in a resume would be a red flag to me if I were hiring. It would come across as entitled.

Re: Atlassian: We estimate the rebuilding effort to last for up to 2 more weeks

#212
post #126
post #92

Earlier quoted context omitted.

I would have just quit if I had that landed on me.

Honestly it doesn't sound too difficult and like a challenge to script something fun. To me. If you ever find yourself in that situation, email is in my profile and we'll work something out :)

No post body was provided.

Re: Atlassian: We estimate the rebuilding effort to last for up to 2 more weeks

#214

The wildest part about this service outage to me is that Atlassian's stock is up 10% over the past month.

Can you imagine screwing up this bad and customers still not leave?

I'll probably should buy their stock.

Re: Atlassian: We estimate the rebuilding effort to last for up to 2 more weeks

#215
post #140

I use JIRA and confluence every single day, I have to, it is everywhere - but imo it is such a horrific toolset in every way (even before this outage), I can't for the life of me figure out how it got so much market-share.

Because all the others are equally bad, basically. The options for enterprisey issue management are quite slim.

I really like ClickUp. Never had an issue.

Re: Atlassian: We estimate the rebuilding effort to last for up to 2 more weeks

#216
Wow! I didn’t realize the scope and duration of this outage. This must be doing some serious damage to some of their clients (catastrophic if this does impact JIRA, Confluence, and OpsGenie broadly on a company level). Is there any report of approximately how many (or specific) companies have been affected as a result of this?

Re: Atlassian: We estimate the rebuilding effort to last for up to 2 more weeks

#217

As someone who is impacted, this is obviously immensely frustrating. Worse, outside of "we have rebuilt functionality for over 35% of the users", I haven't seen any reports from the people who have ostensibly been recovered. Next, their published RTO is 6 hours, so obviously they must have done something that completely demolished their ability to use their standard recovery methods: https://www.atlassian.com/trust/s…

We are one of the customers that has been restored with access to systems restored about 24 hours ago. We have some lingering but relatively minor issues with some plugins and the like but there doesn't appear to be any data loss and performance is healthy.

Re: Atlassian: We estimate the rebuilding effort to last for up to 2 more weeks

#218
post #8
post #4

For those of us not up to date, what exactly has happened? Their status page hasn't actually shown why they're having to rebuild.

>While running a maintenance script, a small number of sites were disabled unintentionally. https://twitter.com/Atlassian/status/1511870509973090304 Most likely they wiped the data

I think sites are their account management stuff. Sounds like they deleted user accounts and not actual data. Notice that only native products are down and not acquisitions. They probably just haven’t migrated those yet.

Re: Atlassian: We estimate the rebuilding effort to last for up to 2 more weeks

#219
post #178

Earlier quoted context omitted.

For what you suggest some combination of these things should have happened. - Some employee has root access to AWS account and uses it operationally - Given wildcard S3 permissions to an IAM user and allowing delete bucket - Not enabled object versioning - Cross Region replication not enabled - no large bucket protection - don't have basic security monitoring and setup of Cloudtail alerts - have not invested in full…

I haven't ever seen it be as perfectly done as you've described. It is always shades of gray across teams and companies. They have most of what you described, but not uniformly across the company. e.g. versioning may be enabled, but not cross region replication because it is cost prohibitive. Someone runs a job to clean up a bucket that includes deleting old versions. They point it at the wrong bucket or wrong path i…

It is hard to get it perfectly right yes , missing by small margins or even doubled/trippled the declared time would be reasonable if it is just prediction problem.

However going like 100x is not probably cause this is hard to get 100 % accurate it look more likely deleted data as being rumoured and more importantly not actually having functioning backups that were ever tested and manually reconstructing from logs and other sources.

More than just RTO, they are not going to be able to meet RPO objectives for affected customers , depending on how much loss that is going to pretty bad.

Re: Atlassian: We estimate the rebuilding effort to last for up to 2 more weeks

#220
post #22

We're mere weeks aware from migrating to their could platform after the self-hosted rugpull. This really doesn't give me confidence in their ability to not break my stuff.

There are two ways this can go: 1) This outage will get their organization to prioritize work such that it never happens again. 2) This outage is representative of a dysfunctional organization that can't prioritize work correctly. If you've been using Atlassian software for a while and are used to how they prioritize tickets then one of those options seems far more likely than the other.

I can tell you #1 never happens. It will be a temporary effect of the management green lighting the years of neglected maintenance work until everybody forgets about it and it will go back to business as usual until the next incident happens and the cycle repeats.
Post reply on HN