Live data from Hacker News

Atlassian: We estimate the rebuilding effort to last for up to 2 more weeks

twitter.com

171–180 of 260 posts

Re: Atlassian: We estimate the rebuilding effort to last for up to 2 more weeks

#171

Earlier quoted context omitted.

that is almost cartoonishly nightmarish

It happened specifically to 1 type of laptop and we only had about 30 of them. So we pulled all of them out of roulation. Then covid struck, so I reformatted most of them with Debian and we gave them away for home schooling. I wonder if I managed to linuxify some kid in the process.

Thanks to you, plenty of kids now think they live 8,000 years in the future! :)

Re: Atlassian: We estimate the rebuilding effort to last for up to 2 more weeks

#172

Atlassian does not care about individual customers. They are purely driven by numbers. I listened to a presentation by one of their founders a long time ago where he admitted that statistics and number management was part of their DNA. They don’t think about the customer’s name, maybe this is right, maybe not.

That's ever single company under capitalism tho. Anything else is only empty platitudes.

Re: Atlassian: We estimate the rebuilding effort to last for up to 2 more weeks

#174

So, let me get this straight: * It's been deleted for a week already, they estimate they might need two more weeks. Three in total. * They claim to have "extensive backups", and hundreds of engineers working on it. What? How? This simply doesn't go together. Why would restoring from backup take three weeks? Either their backups aren't complete, or they need new software written for the restore, or something else does…

If you would permanently delete data for selected customers from a large multitenant system, it could actually take some time to restore it - even with proper backups.

You can’t just do a full recovery as that would mess those customers who were not affected (it likely takes time to notice the mistake - others have continued to use the system). You might need to write some tools to migrate the data from backups. Also you really need to test everything very carefully - otherwise you might be in even deeper trouble (looking at corrupted instead of lost data).

In large organization this kind of ”manual” recovery might require people from multiple teams as no single person knows all the areas. This adds overhead. Throwing too many people in does not help either. When you start thinking about it, few weeks is not that long.

And JIRA is definitely not simple. It’s complicated beast and likely the SaaS features combined with all the legacy makes it even more complicated.

Re: Atlassian: We estimate the rebuilding effort to last for up to 2 more weeks

#175

I bet they screwed up royally, deleted some data and are down to either rebuilding it from logs, caches or other side-effects, or using data recovery software on the storage drives (which might involve third-party companies). I can't see many other reasons why this should take 2 weeks.

Let me bet: Rebuilding from Jira email notifications. Yes, the diffs in the notifications.

> Yes, the diffs in the notifications.

In my limited experience these diffs can be missing information. I recently had to reconstruct an issue description using these email diffs after two people where editing the description at the same time and it was not 100% accurate, several lines were missing. Going to the 'history' tab on the issue I was able to get the missing lines however, if all you have are emails though you might be out of luck.

Re: Atlassian: We estimate the rebuilding effort to last for up to 2 more weeks

#176
post #124

As someone who is impacted, this is obviously immensely frustrating. Worse, outside of "we have rebuilt functionality for over 35% of the users", I haven't seen any reports from the people who have ostensibly been recovered. Next, their published RTO is 6 hours, so obviously they must have done something that completely demolished their ability to use their standard recovery methods: https://www.atlassian.com/trust/s…

RTO is so hard to state properly even with regular testing. If someone blows away a critical database, sure you can meet your published RTO. What if we lose 300 of our databases and need to copy snapshots from another region. AWS limits you to 20 concurrent snapshot copies cross region. Which of those databases should you do first? Do you know your entire dependency graph for all 1000 of your services to make the rig…

Normally for short RTO you do not try to recover by to the orginal location

you fail over to warm replica;s that are already staged with data with in your RPO.

Re: Atlassian: We estimate the rebuilding effort to last for up to 2 more weeks

#177
post #167

I bet they screwed up royally, deleted some data and are down to either rebuilding it from logs, caches or other side-effects, or using data recovery software on the storage drives (which might involve third-party companies). I can't see many other reasons why this should take 2 weeks.

this is a scary thought. I need to start being more aggressive about backing things up that are "in the cloud."

The cloud is just someone else's servers.

3-2-1-0 Applies to all data, at all time, in all places

Re: Atlassian: We estimate the rebuilding effort to last for up to 2 more weeks

#178
post #124

Earlier quoted context omitted.

RTO is so hard to state properly even with regular testing. If someone blows away a critical database, sure you can meet your published RTO. What if we lose 300 of our databases and need to copy snapshots from another region. AWS limits you to 20 concurrent snapshot copies cross region. Which of those databases should you do first? Do you know your entire dependency graph for all 1000 of your services to make the rig…

For what you suggest some combination of these things should have happened. - Some employee has root access to AWS account and uses it operationally - Given wildcard S3 permissions to an IAM user and allowing delete bucket - Not enabled object versioning - Cross Region replication not enabled - no large bucket protection - don't have basic security monitoring and setup of Cloudtail alerts - have not invested in full…

I haven't ever seen it be as perfectly done as you've described. It is always shades of gray across teams and companies. They have most of what you described, but not uniformly across the company.

e.g. versioning may be enabled, but not cross region replication because it is cost prohibitive. Someone runs a job to clean up a bucket that includes deleting old versions. They point it at the wrong bucket or wrong path in the bucket. Or a malicious user does it on purpose. Monitors and alerts really tell you after the fact that you now have a major problem.

Also limits (like cross region concurrency) may not be known about until it is time to actually do a mass scale restore. DR tests might have been done but only in isolation of one app at a time. By the time you realize your mistake you're dealing with physics. Maybe AWS can bump it a bit to help you in that particular circumstance though.

No idea what happened at Atlassian. My only point is it is very hard to get it right without a huge amount of effort.

Re: Atlassian: We estimate the rebuilding effort to last for up to 2 more weeks

#179
post #162

Earlier quoted context omitted.

> Jira Cloud is incredibly slow I haven’t used Jira Cloud in any great depth, but I did play around with it as part of trialling the free plan and was amused that you could quite easily come across warnings to backup your installation and consult your system administrator before proceeding…how exactly do I do that for a cloud service? Doesn’t exactly inspire confidence. At my last job we used Bitbucket Cloud and that…

> I haven’t used Jira Cloud in any great depth. I have. It's painfully slow. But slowness isn't really the problem, the problem is that it's unpredictable. I wait for the interface to be fully loaded, so I click on a text box and i start typing. Then FU--ING something takes the focus to some other element in the web page and now i'm typing random shortcuts (like reassigning tickets, changing status or whatever). It's…

Well you’re just a user, what do you know anyhow? It’s not like paying money for a service entitles you to be able to have something which works and delivers what you paid for. In fact, if you look at the EULA I’m positive that it states you’re paying to access the almighty godlike code of Jira. Also if you don’t like it, simply construct your own industry standard and train your users on it, maintain it, and keep the costs down!.

I think software companies need to have a serious “Come to Jesus” talk with their users about who needs to control what.

Re: Atlassian: We estimate the rebuilding effort to last for up to 2 more weeks

#180

Looking at the company financials and timing, the wooden headquarters might have to wait. I don't envy engineers there right now, some dpts. Stay strong, don't burn out! https://architectureau.com/articles/worlds-tallest-hybrid-ti...

I bet it does not even blimp their financials

Sadly it seems society has assimilated vendors behaving poorly. Data breaches and incompetence does not seem to phase purchasing choices these days

Post reply on HN