Live data from Hacker News

Slack AI Training with Customer Data

slack.com

111–120 of 426 posts

Re: Slack AI Training with Customer Data

#111
post #89
post #61

> We offer Customers a choice around these practices. If you want to exclude your Customer Data from helping train Slack global models, you can opt out. If you opt out, Customer Data on your workspace will only be used to improve the experience on your own workspace and you will still enjoy all of the benefits of our globally trained AI/ML models without contributing to the underlying models. Why would anyone not opt…

> Why would anyone not opt-out? Because you might actually want to have the best possible global models ? Think of "not opting out" as "helping them build a better product". You are already paying for that product, if there is anything you can do, for free and without any additional time investment on your side that makes their next release better, why not do it ? You gain a better product for the same price, they ge…

> Think of "not opting out" as "helping them build a better product"

I feel like someone would only have this opinion if they've never ever dealt with any in the tech industry, or capitalist, in their entire life. So like 8-19 year olds? Except even they seem to understand that the profit absolutist goals undermine everything.

This idea has the same smell as "We're a family" company meetings.

Re: Slack AI Training with Customer Data

#113
post #79

Earlier quoted context omitted.

LLMs are state-of-the-art at detecting sarcasm. It won't help if the data is just wrong though. Edit: https://arxiv.org/abs/2312.03706 Human performance on this benchmark (detecting sarcasm in Reddit comments) was 0.82, a BERT-based LLM scored 0.79. https://arxiv.org/abd/2106.05752 LSTM, 98% at detecting sarcasm in a Project Gutenburg-based dataset.

I literally can't tell if you're being sarcastic or not.

Exactly

Re: Slack AI Training with Customer Data

#114

Story time. I was at a VC conference last year and if I learned nothing else there, I learned how to spell "AI". Every single exhibitor just about had their signage proudly proclaiming their capabilities in this area, but one in particular struck me. They were touting the API integrations they could offer to train their "Enterprise AI"/LLM, and among those integrations were things like M365, Slack, etc. It struck me…

Same for Reddit or Facebook groups. There's a lot of shitposting there, but absolutely a lot of valuable information if LLMs manage to separate the wheat from the chaff.

Re: Slack AI Training with Customer Data

#115
The gold rush for data is wild. Private companies selling us out.

- Slack

- Discord

- Reddit

- Stackoverflow

Let’s just hope this data gold rush dies out faster than the web3 craze before OpenAI reaches critical mass and gets access to government server farms.

Alphabet boys have server farms of domestic and foreign surveillance and intelligence. Exabytes of data [1]

[1] https://en.m.wikipedia.org/wiki/Utah_Data_Center

Re: Slack AI Training with Customer Data

#116
> we do not build or train these models in such a way that they could learn, memorise, or be able to reproduce some part of Customer Data

They don't "build" them this way (whatever that means) but if training data is somehow leaked, they're off the hook because they didn't build it that way?

Re: Slack AI Training with Customer Data

#117
> Contact us to opt out. If you want to exclude your Customer Data from Slack global models, you can opt out. To opt out, please have your Org or Workspace Owners or Primary Owner contact our Customer Experience team at feedback@slack.com with your Workspace/Org URL and the subject line “Slack Global model opt-out request.” We will process your request and respond once the opt out has been completed.

This is not ok. We didn't have to reach out by email to sign up, this should be a toggle in the UI. This is deliberately high friction.

Re: Slack AI Training with Customer Data

#118

Story time. I was at a VC conference last year and if I learned nothing else there, I learned how to spell "AI". Every single exhibitor just about had their signage proudly proclaiming their capabilities in this area, but one in particular struck me. They were touting the API integrations they could offer to train their "Enterprise AI"/LLM, and among those integrations were things like M365, Slack, etc. It struck me…

The best LLMs were trained on data from the open internet, which is full of garbage. They still do a pretty good job (granted it has been fine tuned and RLHF'd, but you can do that with Slack data too)

Re: Slack AI Training with Customer Data

#119
post #61

> We offer Customers a choice around these practices. If you want to exclude your Customer Data from helping train Slack global models, you can opt out. If you opt out, Customer Data on your workspace will only be used to improve the experience on your own workspace and you will still enjoy all of the benefits of our globally trained AI/ML models without contributing to the underlying models. Why would anyone not opt…

Because it's default opt-in, and most people won't see this announcement.

Re: Slack AI Training with Customer Data

#120
post #79

Earlier quoted context omitted.

This assumes everything in the training data set is accurate. Sometimes people are wrong, obtuse, sarcastic, etc. LLM's don't have any way of detecting or accounting for this, do they? That output, then being used to train other LLM's, just creates an ouroboros of AI generated dogshit.

LLMs are state-of-the-art at detecting sarcasm. It won't help if the data is just wrong though. Edit: https://arxiv.org/abs/2312.03706 Human performance on this benchmark (detecting sarcasm in Reddit comments) was 0.82, a BERT-based LLM scored 0.79. https://arxiv.org/abd/2106.05752 LSTM, 98% at detecting sarcasm in a Project Gutenburg-based dataset.

> LLMs are state-of-the-art at detecting sarcasm.

This is such a precious gem.

Post reply on HN