Live data from Hacker News

Slack AI Training with Customer Data

slack.com

361–370 of 426 posts

Re: Slack AI Training with Customer Data

#361

They are breaking a lot of laws by doing this. There are plenty of regulated industries in healthcare and financial services using slack.

Indeed. I flagged this internally the first thing this morning, and fully expect to field questions from our clients over the next couple of weeks as the news percolates through to upper echelons.

I can tolerate Slack and/or Salesforce at large building per-customer overlays on top of a generic LLM. Those at least can provide actual business value[ß], and give their AI teams something to experiment on. But feeding gazillion companies' internal (and workspace-joined!) chats to a global model? Hell no.

Unsurprisingly, we opted out a few hours ago.

ß: a smart, context-aware autocomplete for those who need to type a lot on their phones would not be a bad idea. The current generation of autocorrupt is obnoxious.

Re: Slack AI Training with Customer Data

#362
post #186

Earlier quoted context omitted.

I'm imagining a corporate slack, with information discussed in channels or private chats that exists nowhere else on the internet.. gets rolled into a model. Then, someone asks a very specific question.. conversationally.. about such a very specific scenario.. Seems plausible confidential data would get out, even if it wasn't attributed to the client. Not that it’s possible to ask an llm how a specific or random comp…

exactly. a fun game to see why it is so hard to prevent this https://gandalf.lakera.ai/

Sometimes the obvious questions are met with a lot of silence.

I don't think I can be the only one who has had a conversation with GPT about something obscure they might know but there isn't much about online, and it either can't find anything... or finds it, and more.

Re: Slack AI Training with Customer Data

#363
post #158

Earlier quoted context omitted.

To add on to this: I think it should be mentioned that Slack says they'll prevent data leakage across workspaces in their model, but don't explain how they do this. They don't seem to go into any detail about their data safeguards and how they're excluding sensitive info from training. Textual is good for this purpose since it redacts PII thus preventing it from being leaked by the trained model. Disclaimer: I work a…

How do you handle proprietary data being leaked? Sure you can easily detect and redact names and phone numbers and addresses, but without significant context it seems difficult to detect whether "11 spices - mix with 2 cups of white flour ... 2/3 teaspoons of salt, 1/2 teaspoons of thyme [...]" is just a normal public recipe or a trade secret kept closely guarded for 70 years

Fair question, but you have to consider the realistic alternatives. For most of our customers inaction isn't an option. The combination of NER models + synthesis LLMs actually handles these types of cases fairly well. I put your comment into our web app and this was the output:

How do you handle proprietary data being leaked? Sure you can easily detect and redact names and phone numbers and addresses, but without significant context it seems difficult to detect whether "17 spices - mix with 2lbs of white flour ... half teaspoon of salt, 1 tablespoon of thyme [...]" is just a normal public recipe or a trade secret kept closely guarded for 75 years.

Re: Slack AI Training with Customer Data

#364
post #252

We really need to start using self-hosted solutions. Like matrix / element for team messaging. It's ok not wanting to run your own hardware at your own premises. But the solution is to run a solution that is end-to-end encrypted so that the hosting service cannot get at the data. cryptpad.fr is another great piece of software.

You could checkout [Campfire](https://once.com/campfire). You get teh source code (Ruby on Rails) and you deploy wherever. We're running ours on a digital ocean droplet.

Re: Slack AI Training with Customer Data

#366
post #147

Earlier quoted context omitted.

Especially when a few paragraphs below they say: > If you want to exclude your Customer Data from helping train Slack global models, you can opt out. So Customer Data is not used to train models "used broadly across all of our customers [in such a way that ...]", but... it is used to help train global models. Uh.

Why are these kinda things opt-out? And need to be discovered.. We're literally discussing switching to Teams at my company (1500 employees)

Possible alternatives that may not be OpenAI behind the scenes.

https://www.producthunt.com/categories/team-collaboration

There's also a Lot of "Best Slack Alternatives in 2024" and such.

Re: Slack AI Training with Customer Data

#367
It’s very difficult to ensure no data leakage, even for something like an Emoji prediction model. If you can try a large number of inputs and see what the suggested emoji is, that’s going to give you information about the training set. I wouldn’t be surprised to see a trading company or two pop up trying to exploit this to get insider information.

Re: Slack AI Training with Customer Data

#370

Earlier quoted context omitted.

Cool. Now do it for a 3,000 person org with varied tech skills and patience.

There is no way a chat room with 3,000 people is remotely useful to anybody.

Right, that's why thousands of enterprises use them.
Post reply on HN