Live data from Hacker News

Slack AI Training with Customer Data

slack.com

271–280 of 426 posts

Re: Slack AI Training with Customer Data

#271

Earlier quoted context omitted.

Personally, I rather liked self-hosted versions of these: Mattermost: https://mattermost.com/ Rocket.Chat: https://www.rocket.chat/ Nextcloud Talk: https://nextcloud.com/talk/ Out of those, Mattermost was the easiest to setup (just need PostgreSQL and a web server, in addition to the main container), however not being able to easily permanently delete instead of just archiving workspaces was awkward. Nextcloud Talk w…

The problem is not technical, but social with these platforms. i.e. How do you convince 40+ people from 5 countries to add yet another memory resident chat application and fragment their knowledge to another app/mental space? This gets way harder as the community becomes more dynamic and temporary (i.e. high circulation like students). I gave the good fight last year with someone, and they just didn't flex a nanomete…

> i.e. How do you convince 40+ people from 5 countries to add yet another memory resident chat application and fragment their knowledge to another app/mental space?

If it's a company, you can just be like: "Hey, we use this platform for communication, you can log in with your Active Directory credentials."

It also has the added benefit of acting as a directory for every employee in the company, so getting in touch can be more convenient than e-mail (while you can also customize the notification preferences, so it doesn't get too spammy), as opposed to the situation which might develop, where some teams or org units are on Slack, others on Teams and getting in touch can be more messy.

If it's a free-form social group, then you can throw that idea away because of network effects, it'd be an uphill battle, same as how sometimes people complain about people using Discord for various communities, but at the same time the reality is that old school forums and such were also killed off - since most people already have a Discord account and there's less friction to just use that.

Either way, I'm happy that self-hosted software like that exists.

Re: Slack AI Training with Customer Data

#272
post #61

> We offer Customers a choice around these practices. If you want to exclude your Customer Data from helping train Slack global models, you can opt out. If you opt out, Customer Data on your workspace will only be used to improve the experience on your own workspace and you will still enjoy all of the benefits of our globally trained AI/ML models without contributing to the underlying models. Why would anyone not opt…

I'm willing to bet that for smaller companies, they just won't care enough to consider this an issue and that's what Slack/Salesforce is hedging on.

I can't see a universe in which large corpos would allow such blatant corporate espionage for a product they pay for no less. But I can already imagine trying to talk my CTO (who is deep into the AI sycophancy) into opting us out is gonna be arduous at best.

Re: Slack AI Training with Customer Data

#273
post #256

Earlier quoted context omitted.

To me it says that they _do_ train global models with customer data, but they are trying to ensure no data leakage (which will be hard, but maybe not impossible, if they are training with it). The caveats are for “local” models, where you would want the model to be able to answer questions about discussions in the workspace. It makes me wonder how they handle “private” chats, can they leak across a workspace? Presuma…

My intuition is that it's impossible to guarantee there are no leaks in the LLM as it stands today. It would surely require some new computer science to ensure that no part of any output that could ever possibly be developed isn't sensitive data from any of the input. It's one thing if the input is the published internet (even if covered by copyright), it's entirely another to be using private training data from corp…

There is a way. Build a preference model from the sensitive dataset. Then use the preference model with RLAIF (like RLHF but with AI instead of humans) to fine-tune the LLM. This way only judgements about the LLM outputs will pass from the sensitive dataset. Copy the sense of what is good, not the data.

Re: Slack AI Training with Customer Data

#274

Earlier quoted context omitted.

The problem is not technical, but social with these platforms. i.e. How do you convince 40+ people from 5 countries to add yet another memory resident chat application and fragment their knowledge to another app/mental space? This gets way harder as the community becomes more dynamic and temporary (i.e. high circulation like students). I gave the good fight last year with someone, and they just didn't flex a nanomete…

> i.e. How do you convince 40+ people from 5 countries to add yet another memory resident chat application and fragment their knowledge to another app/mental space? If it's a company, you can just be like: "Hey, we use this platform for communication, you can log in with your Active Directory credentials." It also has the added benefit of acting as a directory for every employee in the company, so getting in touch ca…

> If it's a company

That's a big if, and the answer is "No" in my case. If it was, that comment wouldn't be there.

It's not a "social group" either, but a group of independent institutions working together. It's like a large gear-train. A lot of connections between small islands of people. So you have to work together, and have to find a way somehow. So, it's complicated.

> Either way, I'm happy that self-hosted software like that exists.

Me too. I happen to manage a Nextcloud instance, but nobody is interested in the "Talk" module.

Re: Slack AI Training with Customer Data

#276

Good we moved to matrix already. I just hope they start putting more emphasis on Element X, which message handling is broken on iOS for weeks now.

Element X is where all the effort is going, and should be working really well. How is msg handling broken?

Not the OP here, but I've tried really hard to use Element X and it crashes constantly.

Re: Slack AI Training with Customer Data

#277

In case this is helpful to anyone else, I opted out earlier today with an email to feedback@slack.com Subject: Slack Global Model opt-out request. Body: .slack.com Please opt the above Slack Workspace out of training of Slack Global Models.

Make sure you put a period at the end of the subject line. Their quoted text includes a period at the end. Please also scold them for behaving unethically and perhaps breaking the law.

The period is outside the quotes though, are you suggesting we should have the quotes too?

Re: Slack AI Training with Customer Data

#279
post #227

Well I really hope this massively blows up in their face when all of Europe goes to work just about now, and then North America in 5-8 hours. Let's see if we have another Helldivers 2 event that makes them do a hard backpedal after losing thousands of large customers that will not under any circumstances take the chance. I have a friend with a law firm who just called me yesterday for advice as he's thinking about sw…

Yeah my lawyer friend is worried he might even lose his license over this. It's gonna be very interesting seeing how legal departments react to this.

If you disagree with practices like this, mention this to your legal.

Re: Slack AI Training with Customer Data

#280
post #20
post #6

I'm confused about this statement: "When developing AI/ML models or otherwise analyzing Customer Data, Slack can’t access the underlying content. We have various technical measures preventing this from occurring" "Can't" is a strong word. I'm curious how an AI model could access data, but Slack, Inc itself couldn't. I suspect they mean "doesn't" instead of "can't", unless I'm missing something.

I also find the word "Slack" in that interesting. I assume they mean "employees of Slack", but the word "Slack" obviously means all the company's assets and agents, systems, computers, servers, AI models, etc. I would find even a statement from Signal like "we can't access our users content" to be tenuous and overly-optimistic. Like, when I heard the word "can't" my brain goes to: there is nothing anyone in the compa…

> I would find even a statement from Signal like "we can't access our users content" to be tenuous and overly-optimistic.

I don't really agree with this statement. Signal literally can't read user data right now. The statement is true, why can't they use it?

If they can't use it, nobody can. there are no services that can't publish an update reversing any security measure available. Also doing that would be illegal, because it would render the statement "we can't access our users content" false.

In Slack case, it is totally different. Data is accessible by Slack systems, the statement "we can't access our users content" is already false. Probably what they mean is something along the lines of: "The data can't be accessed by our systems, but we have measures in place that block the access to most of our employees"

Post reply on HN