Live data from Hacker News

Atlassian enables default data collection to train AI

letsdatascience.com

141–149 of 149 posts

Re: Atlassian enables default data collection to train AI

#141

Atlassian just goes from misstep to misstep. I still use their products quite often. The amount of P0 bugs I experience is absolutely crazy: - Bitbucket workers are hopelessly out of date (self hosted). We've had to put so many random workarounds in especially for Docker, as they don't keep them up to date enough - I have had a bug in JIRA for years where I can't reorder a new ticket unless I refresh the page - Every…

> Anyone have any insight into why things have got so so dysfunctional?

My theory is that there is no incentive for things to not be dysfunctional. At my org anyway Atlassian is almost as entrenched as Google, and we can't even conceive of using something different.

I use plane.so for personal projects after learning about it here. It's a lot simpler and the UI takes some getting used to, but it's fast and it works.

Re: Atlassian enables default data collection to train AI

#142
post #118

Yet another opportunity to provide an alternative that keeps data private

I’m building a self-hosted Confluence alternative called Docmost. It’s open-source and can run fully air-gapped. GitHub: https://github.com/docmost/docmost

[dead]

Re: Atlassian enables default data collection to train AI

#143

Earlier quoted context omitted.

Steam seems to be doing pretty well after 20 years

Valve isn't a publicly traded company with shareholders looking to extract every possible nickel out of the company at any cost.

Think we've found the solution right there

Re: Atlassian enables default data collection to train AI

#144

Plenty of other companies enable this by default too, such as Github, Figma, Adobe, Vercel. I think it's fair to assume that if you ahve data stored within any company, they'll by default use it for training.

Do you have any source for Github training on private repositories if you don't interact with copilot yourself?

I believe you have to disable the toggle, but not sure if it applies to folks who dont use copilot: https://github.com/orgs/community/discussions/188488

Re: Atlassian enables default data collection to train AI

#145
post #44
post #2

I really wish I could find a better source to link to for this. By default, all free and paid customers are being opted-in to their data being used for AI training. All your Confluence pages, Jira tickets, etc. https://support.atlassian.com/security-and-access-policies/d... describes how to disable this, but it also appears that the setting to disable this doesn't exist (it's not visible on any of our instances).

What about really sensitive stuff like if possibly private tickets that have all kinds of stuff like customer data, embargoed CVE fixes or even sensitive health related data, are they just cobble that all into a model so it can leak out to random people ?

> are they just cobble that all into a model so it can leak out to random people ?

Yes. Industrial espionage made easy.

Re: Atlassian enables default data collection to train AI

#147
A to my knowledge unconfirmed claim by the register states:

"Tseytlin said that some Atlassian customers are completely excluded from metadata or in‑app "data contribution" entirely. This includes those who use customer-managed keys, or bring your own key, Atlassian Government Cloud, or Atlassian Isolated Cloud users. He said Atlassian will also not collect metadata or in-app data from customers with HIPAA compliance requirements or from some government and financial services customers."

https://www.theregister.com/2026/04/18/atlassians_new_data_c...

Re: Atlassian enables default data collection to train AI

#148
post #120

You can thank GitHub for setting this draconian precedent

I don't think GitHub even set a precedent for this. My understanding is that they don't train on private repositories per se, though if you access a private repository through copilot, the data flow through copilot can be trained on, which pulls in data from the repo. So a private repo should be safe, as long as you don't use copilot. While Atlassian wants to pull in data from private issue trackers/wikis.

Listen to yourself. Take a moment and try to unpack the mental gymnastics wrangling you just did. Ask yourself, why does the fact that you have a Copilot subscription make it okay to train on all your private repos?

GitHub does not have any of its own models. It routes to partners like OpenAI. Just because some data is from private repos, doesn’t mean all data is flowing nor does it mean it should be trained on just because it’s being inferences on, and there is a difference on the data that was used vs. all the data from that repo, and difference between just that repo vs. all private repos. And they made it all opted in as default. Draconian.

So yes, they did set a precedent and you’re here arguing why it’s okay.

Re: Atlassian enables default data collection to train AI

#149

Earlier quoted context omitted.

Steam seems to be doing pretty well after 20 years

Valve isn't a publicly traded company with shareholders looking to extract every possible nickel out of the company at any cost.

But they are not that much better. We are talking about the company that single-handedly normalized not owning your games, not to mention their contributions to microtransactions or getting children into gambling.
Post reply on HN