Live data from Hacker News

Slack AI Training with Customer Data

slack.com

71–80 of 426 posts

Re: Slack AI Training with Customer Data

#71
post #61

> We offer Customers a choice around these practices. If you want to exclude your Customer Data from helping train Slack global models, you can opt out. If you opt out, Customer Data on your workspace will only be used to improve the experience on your own workspace and you will still enjoy all of the benefits of our globally trained AI/ML models without contributing to the underlying models. Why would anyone not opt…

> Why would anyone not opt-out?

This is basically like all privacy on the internet.

Everyone WOULD opt-out, if it was easy, and it becomes a whack-a-optput game.

note how you opt-out (generic contact us), and what happens when you do opt-out (they still train anyway)

Re: Slack AI Training with Customer Data

#72

>Our mission is to build a product that makes work life simpler, more pleasant and more productive. I know it would be impossible but I wish we go back to the days when we didn't have Slack (or tools alike). Our Slack is a cesspool of people complaining, talking behind other people's backs, echo chamber of negativity etc. That probably speaks more to the overall culture of the company, but Slack certainly doesn't hel…

I disagree slack plays a role. You only mentioned human aspects, nothing to do with technology. There was always going to be instant messaging as software once computers and networks were invented. You'd just say this happens over email and blame email.

Re: Slack AI Training with Customer Data

#73

In case this is helpful to anyone else, I opted out earlier today with an email to feedback@slack.com Subject: Slack Global Model opt-out request. Body: .slack.com Please opt the above Slack Workspace out of training of Slack Global Models.

Make sure you put a period at the end of the subject line. Their quoted text includes a period at the end. Please also scold them for behaving unethically and perhaps breaking the law.

We just opted out. I told them our lawyers have been instructed to watch them like a hawk.

Re: Slack AI Training with Customer Data

#74
In summary, you must opt-out if you want to exclude your data from global models.

Incredibly confusing language since they also vaguely state that "data will not leak across workspaces".

Use tools that cannot leak data not "will not".

Re: Slack AI Training with Customer Data

#77
post #61

> We offer Customers a choice around these practices. If you want to exclude your Customer Data from helping train Slack global models, you can opt out. If you opt out, Customer Data on your workspace will only be used to improve the experience on your own workspace and you will still enjoy all of the benefits of our globally trained AI/ML models without contributing to the underlying models. Why would anyone not opt…

> We offer Customers a choice around these practices.

I remembered the joke from The Hitchhiker's Guide to the Galaxy, maybe they will have a small hint in a very inconspicuous place, like inserting this into the user agreement on page 300 or so.

Re: Slack AI Training with Customer Data

#78
post #37

Earlier quoted context omitted.

> Data will not leak across workspaces. For any model that will be used broadly across all of our customers, we do not build or train these models in such a way that they could learn, memorize, or be able to reproduce some part of Customer Data. I feel like that is explicitly what this is saying.

The problem is, it's really really hard to guarantee that. Yes if they only train say, classifiers, then the only thing that can leak is the classification outcome. But these things can be super subtle. Even a classifier could leak things if you can hack the context fed into it. They are really playing with fire here.

yes, i certainly agree with you. i think oftentimes these policies are written by non-technical people

i'm not entirely convinced that classifiers and LLMs are disjoint to begin with

Re: Slack AI Training with Customer Data

#79
post #51

Earlier quoted context omitted.

Yes, consider an existing LLM being given “shitpost-y” messages and asking it if there is anything interesting in there. It could probably summarize it well and that could then be used for training another LLM. etc etc

This assumes everything in the training data set is accurate. Sometimes people are wrong, obtuse, sarcastic, etc. LLM's don't have any way of detecting or accounting for this, do they? That output, then being used to train other LLM's, just creates an ouroboros of AI generated dogshit.

LLMs are state-of-the-art at detecting sarcasm. It won't help if the data is just wrong though.

Edit: https://arxiv.org/abs/2312.03706 Human performance on this benchmark (detecting sarcasm in Reddit comments) was 0.82, a BERT-based LLM scored 0.79.

https://arxiv.org/abd/2106.05752 LSTM, 98% at detecting sarcasm in a Project Gutenburg-based dataset.

Re: Slack AI Training with Customer Data

#80
post #20
post #6

I'm confused about this statement: "When developing AI/ML models or otherwise analyzing Customer Data, Slack can’t access the underlying content. We have various technical measures preventing this from occurring" "Can't" is a strong word. I'm curious how an AI model could access data, but Slack, Inc itself couldn't. I suspect they mean "doesn't" instead of "can't", unless I'm missing something.

I also find the word "Slack" in that interesting. I assume they mean "employees of Slack", but the word "Slack" obviously means all the company's assets and agents, systems, computers, servers, AI models, etc. I would find even a statement from Signal like "we can't access our users content" to be tenuous and overly-optimistic. Like, when I heard the word "can't" my brain goes to: there is nothing anyone in the compa…

Interestingly I'd probably go the other way.

If it's verifiably E2EE then I consider "we can't access this" to be a fairly powerful statement. Sure, the source could change, but if you have a reasonable distribution mechanism (e.g. all users get the same code, verifiably reproducible) then that's about as good as you can get.

Privacy policies that state "we won't do XYZ" have literally zero value to me to the extent that I don't even look at them. If I give you some data, it's already leaked in my mind, it's just a matter of time.

Post reply on HN