Live data from Hacker News

Slack AI Training with Customer Data

slack.com

141–150 of 426 posts

Re: Slack AI Training with Customer Data

#141
post #121
post #74

In summary, you must opt-out if you want to exclude your data from global models. Incredibly confusing language since they also vaguely state that "data will not leak across workspaces". Use tools that cannot leak data not "will not".

Most, if not all SaaS software is multi-tenant, so we've been living in the "will not" world for decades now.

In your experience.

Re: Slack AI Training with Customer Data

#142
post #61

> We offer Customers a choice around these practices. If you want to exclude your Customer Data from helping train Slack global models, you can opt out. If you opt out, Customer Data on your workspace will only be used to improve the experience on your own workspace and you will still enjoy all of the benefits of our globally trained AI/ML models without contributing to the underlying models. Why would anyone not opt…

> We offer Customers a choice around these practices. I remembered the joke from The Hitchhiker's Guide to the Galaxy, maybe they will have a small hint in a very inconspicuous place, like inserting this into the user agreement on page 300 or so.

Never more true than with apple.

Activating an iphone for example has a screen devoted to how privacy is important!

It will show you literally thousands of pages of how they take privacy seriously!

(and you can't say NO anywhere in the dialog, they just show you)

They are normalizing "you cannot do anything", and then everyone does it.

Re: Slack AI Training with Customer Data

#144
I wonder how many people that are really mad about these guys or SE using their professional output to train models thought commercial artists were just being whiny sore losers when Deviant Art, Adobe, OpenAI, Stability, et al did it to them.

Re: Slack AI Training with Customer Data

#146
post #74

In summary, you must opt-out if you want to exclude your data from global models. Incredibly confusing language since they also vaguely state that "data will not leak across workspaces". Use tools that cannot leak data not "will not".

Aren't all tools, essentially, just one API call away from "leaking data"?

In this case they mean leak into the global model — so no. You can have sovereignty of your data if you use an open protocol like IRC or Matrix, or a self-hosted tool like Zulip, Mattermost, Rocket Chat, etc

Re: Slack AI Training with Customer Data

#147
post #43

> For any model that will be used broadly across all of our customers, we do not build or train these models in such a way that they could learn, memorise, or be able to reproduce some part of Customer Data This feels so full of subtle qualifiers and weasel words that it generates far more distrust than trust. It only refers to models used "broadly across all" customers - so if it's (a) not used "broadly" or (b) only…

Especially when a few paragraphs below they say:

> If you want to exclude your Customer Data from helping train Slack global models, you can opt out.

So Customer Data is not used to train models "used broadly across all of our customers [in such a way that ...]", but... it is used to help train global models. Uh.

Re: Slack AI Training with Customer Data

#148

Story time. I was at a VC conference last year and if I learned nothing else there, I learned how to spell "AI". Every single exhibitor just about had their signage proudly proclaiming their capabilities in this area, but one in particular struck me. They were touting the API integrations they could offer to train their "Enterprise AI"/LLM, and among those integrations were things like M365, Slack, etc. It struck me…

Most shitposting is probably more straightforward to understand than business communication or press releases where realizing what wasn't said often carries more insight than the things that were said.

Of course training an AI model on simple, straightforward and honest data provides good results. That's the essence behind "textbooks is all you need" which lead to the phi LLMs. Those are great small LLMs. But if you want your model to understand the complexity of human communication you have to include it in your training data.

If you subscribe to the idea that to be the very best text completion engine possible you would need to have a perfect understanding of reality itself, how different humans perceive reality differently, and how they choose to communicate about this perception and their interaction with reality, themselves and other humans, then it's not unreasonable to expect that back-propagation would eventually find that optimal representation if given enough data, the right architecture and enough processing power. Or at least come somewhat close. In that paradigm there is no "bad data", only insufficient or badly balanced datasets. Just don't try doing that with a 3B parameter LLM.

Re: Slack AI Training with Customer Data

#149
post #121
post #74

In summary, you must opt-out if you want to exclude your data from global models. Incredibly confusing language since they also vaguely state that "data will not leak across workspaces". Use tools that cannot leak data not "will not".

Most, if not all SaaS software is multi-tenant, so we've been living in the "will not" world for decades now.

That's exactly my point. "File over app"[1] is just as relevant for businesses as it is for individuals — if you don't want your data to be used for training, then take sovereignty of it.

[1] https://stephango.com/file-over-app

Post reply on HN