Live data from Hacker News

Slack AI Training with Customer Data

slack.com

181–190 of 426 posts

Re: Slack AI Training with Customer Data

#181
post #172

Earlier quoted context omitted.

What do you think chatGPT uses as training data? The whole world’s “sh*tposting”: Reddit, blogs, and the rest of the internet. But also books and Wikipedia and what not. You can “smooth” all the crap out via the training procedure. But even more, Slack can easily filter training data to, say, only posts in high-use channels. Further, slack has other options: eg, use their customer data only for marginal fine-tuning,…

What makes you think I don't shitpost in the #engineering channel? And heuristics don't even scratch the surface of the bigger problem where it's trained on people who aren't great at their jobs but type a lot of words on slack about circling back on KPIs.

I think those types of people are actually shockingly well paid. If slack can make bots to replace them, they’ll print money, right?

Re: Slack AI Training with Customer Data

#182

>Our mission is to build a product that makes work life simpler, more pleasant and more productive. I know it would be impossible but I wish we go back to the days when we didn't have Slack (or tools alike). Our Slack is a cesspool of people complaining, talking behind other people's backs, echo chamber of negativity etc. That probably speaks more to the overall culture of the company, but Slack certainly doesn't hel…

Your company sucks. I’ve used slack at four workplaces and it’s not been at all like that. A previous company had mailing lists and they were toxic as you describe. The tool was not the issue.

Yeah, written communication is harder than in-person communication.

It’s easy to come across poorly in writing, but that issue has no easy resolution unless you’re prepared to ban Slack, email, and any other text-based communication system between employees.

Slack can sometimes be a place for people who don’t feel heard in conventional spaces to vent — but that’s an organisational problem, not a Slack problem.

Re: Slack AI Training with Customer Data

#184
post #51

Earlier quoted context omitted.

Yes, consider an existing LLM being given “shitpost-y” messages and asking it if there is anything interesting in there. It could probably summarize it well and that could then be used for training another LLM. etc etc

This assumes everything in the training data set is accurate. Sometimes people are wrong, obtuse, sarcastic, etc. LLM's don't have any way of detecting or accounting for this, do they? That output, then being used to train other LLM's, just creates an ouroboros of AI generated dogshit.

And yet human civilization has survived the fact that many humans are wrong, lying, delusional, etc. There is no assumption that everything in our personal training set is accurate. In fact, things work better when we explicitly reject that idea.

LLMs do not rely on 100% factually accurate inputs. Sure, you’d rather have less BS than more, but this is all statistics. Just like most people realize that flat earthers are nutty, LLMs can ingest falsehoods without reducing output quality (again, subject to statistics)

Re: Slack AI Training with Customer Data

#185
> Contact us to opt out. If you want to exclude your Customer Data from Slack global models, you can opt out. To opt out, please have your org, workspace owners or primary owner contact our Customer Experience team at feedback@slack.com

Sounds like an invitation for malicious compliance. Anyone can email them a huge text with workspace buried somewhere and they have to decipher it somehow.

Example [Answer is Org-12-Wp]:

"

FORMAL DIRECTIVE AND BINDING COVENANT

WHEREAS, the Parties to this Formal Directive and Binding Covenant, to wit: [Your Name] (hereinafter referred to as "Principal") and [AI Company Name] (hereinafter referred to as "Technological Partner"), wish to enter into a binding agreement regarding certain parameters for the training of an artificial intelligence system;

AND WHEREAS, the Principal maintains control and discretion over certain proprietary data repositories constituting segmented information habitats;

AND WHEREAS, the Principal desires to exempt one such segmented information habitat, namely the combined loci identified as "Org", the region denoted as "12", and the territory designated "Wp", from inclusion in the training data utilized by the Technological Partner for machine learning purposes;

NOW, THEREFORE, in consideration of the mutual covenants and promises contained herein, the receipt and sufficiency of which are hereby acknowledged, the Parties agree as follows:

DEFINITIONS

1.1 "Restricted Information Habitat" shall refer to the proprietary data repository identified by the Principal as the conjoined loci of "Org", the region "12", and the territory "Wp".

OBLIGATIONS OF TECHNOLOGICAL PARTNER

2.1 The Technological Partner shall implement all reasonably necessary technical and organizational measures to ensure that the Restricted Information Habitat, as defined herein, is excluded from any training data sets utilized for machine learning model development and/or refinement.

2.2 The Technological Partner shall maintain an auditable record of compliance with the provisions of this Formal Directive and Binding Covenant, said record being subject to inspection by the Principal upon reasonable notice.

REMEDIES

3.1 In the event of a material breach...

[Additional legalese]

IN WITNESS WHEREOF, the Parties have executed this Formal Directive and Binding Covenant."

Re: Slack AI Training with Customer Data

#186
post #43

> For any model that will be used broadly across all of our customers, we do not build or train these models in such a way that they could learn, memorise, or be able to reproduce some part of Customer Data This feels so full of subtle qualifiers and weasel words that it generates far more distrust than trust. It only refers to models used "broadly across all" customers - so if it's (a) not used "broadly" or (b) only…

I'm imagining a corporate slack, with information discussed in channels or private chats that exists nowhere else on the internet.. gets rolled into a model.

Then, someone asks a very specific question.. conversationally.. about such a very specific scenario..

Seems plausible confidential data would get out, even if it wasn't attributed to the client.

Not that it’s possible to ask an llm how a specific or random company in an industry might design something…

Re: Slack AI Training with Customer Data

#187
post #147
post #43

> For any model that will be used broadly across all of our customers, we do not build or train these models in such a way that they could learn, memorise, or be able to reproduce some part of Customer Data This feels so full of subtle qualifiers and weasel words that it generates far more distrust than trust. It only refers to models used "broadly across all" customers - so if it's (a) not used "broadly" or (b) only…

Especially when a few paragraphs below they say: > If you want to exclude your Customer Data from helping train Slack global models, you can opt out. So Customer Data is not used to train models "used broadly across all of our customers [in such a way that ...]", but... it is used to help train global models. Uh.

Hope it's not doublespeak, ambiguity leaves it grey, maybe to play.

Re: Slack AI Training with Customer Data

#188
post #147
post #43

> For any model that will be used broadly across all of our customers, we do not build or train these models in such a way that they could learn, memorise, or be able to reproduce some part of Customer Data This feels so full of subtle qualifiers and weasel words that it generates far more distrust than trust. It only refers to models used "broadly across all" customers - so if it's (a) not used "broadly" or (b) only…

Especially when a few paragraphs below they say: > If you want to exclude your Customer Data from helping train Slack global models, you can opt out. So Customer Data is not used to train models "used broadly across all of our customers [in such a way that ...]", but... it is used to help train global models. Uh.

Opt out is such bullshit.

Re: Slack AI Training with Customer Data

#189
post #43

> For any model that will be used broadly across all of our customers, we do not build or train these models in such a way that they could learn, memorise, or be able to reproduce some part of Customer Data This feels so full of subtle qualifiers and weasel words that it generates far more distrust than trust. It only refers to models used "broadly across all" customers - so if it's (a) not used "broadly" or (b) only…

If you trained on customer data your service contains custom data.

Re: Slack AI Training with Customer Data

#190

Earlier quoted context omitted.

Why are these kinda things opt-out? And need to be discovered.. We're literally discussing switching to Teams at my company (1500 employees)

You’d be better off just not having chat than switching to Teams.

But the business will suffer by most likely being less successful due to less cohesive communication.. it's "The Ick" either way.
Post reply on HN