Earlier quoted context omitted.
What do you think chatGPT uses as training data? The whole world’s “sh*tposting”: Reddit, blogs, and the rest of the internet. But also books and Wikipedia and what not. You can “smooth” all the crap out via the training procedure. But even more, Slack can easily filter training data to, say, only posts in high-use channels. Further, slack has other options: eg, use their customer data only for marginal fine-tuning,…
What makes you think I don't shitpost in the #engineering channel? And heuristics don't even scratch the surface of the bigger problem where it's trained on people who aren't great at their jobs but type a lot of words on slack about circling back on KPIs.
Slack AI Training with Customer Data
181–190 of 426 posts
Re: Slack AI Training with Customer Data
#182>Our mission is to build a product that makes work life simpler, more pleasant and more productive. I know it would be impossible but I wish we go back to the days when we didn't have Slack (or tools alike). Our Slack is a cesspool of people complaining, talking behind other people's backs, echo chamber of negativity etc. That probably speaks more to the overall culture of the company, but Slack certainly doesn't hel…
Your company sucks. I’ve used slack at four workplaces and it’s not been at all like that. A previous company had mailing lists and they were toxic as you describe. The tool was not the issue.
It’s easy to come across poorly in writing, but that issue has no easy resolution unless you’re prepared to ban Slack, email, and any other text-based communication system between employees.
Slack can sometimes be a place for people who don’t feel heard in conventional spaces to vent — but that’s an organisational problem, not a Slack problem.
Re: Slack AI Training with Customer Data
#183Re: Slack AI Training with Customer Data
#184Earlier quoted context omitted.
Yes, consider an existing LLM being given “shitpost-y” messages and asking it if there is anything interesting in there. It could probably summarize it well and that could then be used for training another LLM. etc etc
This assumes everything in the training data set is accurate. Sometimes people are wrong, obtuse, sarcastic, etc. LLM's don't have any way of detecting or accounting for this, do they? That output, then being used to train other LLM's, just creates an ouroboros of AI generated dogshit.
LLMs do not rely on 100% factually accurate inputs. Sure, you’d rather have less BS than more, but this is all statistics. Just like most people realize that flat earthers are nutty, LLMs can ingest falsehoods without reducing output quality (again, subject to statistics)
Re: Slack AI Training with Customer Data
#185Sounds like an invitation for malicious compliance. Anyone can email them a huge text with workspace buried somewhere and they have to decipher it somehow.
Example [Answer is Org-12-Wp]:
"
FORMAL DIRECTIVE AND BINDING COVENANT
WHEREAS, the Parties to this Formal Directive and Binding Covenant, to wit: [Your Name] (hereinafter referred to as "Principal") and [AI Company Name] (hereinafter referred to as "Technological Partner"), wish to enter into a binding agreement regarding certain parameters for the training of an artificial intelligence system;
AND WHEREAS, the Principal maintains control and discretion over certain proprietary data repositories constituting segmented information habitats;
AND WHEREAS, the Principal desires to exempt one such segmented information habitat, namely the combined loci identified as "Org", the region denoted as "12", and the territory designated "Wp", from inclusion in the training data utilized by the Technological Partner for machine learning purposes;
NOW, THEREFORE, in consideration of the mutual covenants and promises contained herein, the receipt and sufficiency of which are hereby acknowledged, the Parties agree as follows:
DEFINITIONS
1.1 "Restricted Information Habitat" shall refer to the proprietary data repository identified by the Principal as the conjoined loci of "Org", the region "12", and the territory "Wp".
OBLIGATIONS OF TECHNOLOGICAL PARTNER
2.1 The Technological Partner shall implement all reasonably necessary technical and organizational measures to ensure that the Restricted Information Habitat, as defined herein, is excluded from any training data sets utilized for machine learning model development and/or refinement.
2.2 The Technological Partner shall maintain an auditable record of compliance with the provisions of this Formal Directive and Binding Covenant, said record being subject to inspection by the Principal upon reasonable notice.
REMEDIES
3.1 In the event of a material breach...
[Additional legalese]
IN WITNESS WHEREOF, the Parties have executed this Formal Directive and Binding Covenant."
Re: Slack AI Training with Customer Data
#186> For any model that will be used broadly across all of our customers, we do not build or train these models in such a way that they could learn, memorise, or be able to reproduce some part of Customer Data This feels so full of subtle qualifiers and weasel words that it generates far more distrust than trust. It only refers to models used "broadly across all" customers - so if it's (a) not used "broadly" or (b) only…
Then, someone asks a very specific question.. conversationally.. about such a very specific scenario..
Seems plausible confidential data would get out, even if it wasn't attributed to the client.
Not that it’s possible to ask an llm how a specific or random company in an industry might design something…
Re: Slack AI Training with Customer Data
#187> For any model that will be used broadly across all of our customers, we do not build or train these models in such a way that they could learn, memorise, or be able to reproduce some part of Customer Data This feels so full of subtle qualifiers and weasel words that it generates far more distrust than trust. It only refers to models used "broadly across all" customers - so if it's (a) not used "broadly" or (b) only…
Especially when a few paragraphs below they say: > If you want to exclude your Customer Data from helping train Slack global models, you can opt out. So Customer Data is not used to train models "used broadly across all of our customers [in such a way that ...]", but... it is used to help train global models. Uh.
Re: Slack AI Training with Customer Data
#188> For any model that will be used broadly across all of our customers, we do not build or train these models in such a way that they could learn, memorise, or be able to reproduce some part of Customer Data This feels so full of subtle qualifiers and weasel words that it generates far more distrust than trust. It only refers to models used "broadly across all" customers - so if it's (a) not used "broadly" or (b) only…
Especially when a few paragraphs below they say: > If you want to exclude your Customer Data from helping train Slack global models, you can opt out. So Customer Data is not used to train models "used broadly across all of our customers [in such a way that ...]", but... it is used to help train global models. Uh.
Re: Slack AI Training with Customer Data
#189> For any model that will be used broadly across all of our customers, we do not build or train these models in such a way that they could learn, memorise, or be able to reproduce some part of Customer Data This feels so full of subtle qualifiers and weasel words that it generates far more distrust than trust. It only refers to models used "broadly across all" customers - so if it's (a) not used "broadly" or (b) only…
Re: Slack AI Training with Customer Data
#190Earlier quoted context omitted.
Why are these kinda things opt-out? And need to be discovered.. We're literally discussing switching to Teams at my company (1500 employees)
You’d be better off just not having chat than switching to Teams.