Live data from Hacker News

Slack AI Training with Customer Data

slack.com

101–110 of 426 posts

Re: Slack AI Training with Customer Data

#102
post #56
post #43

> For any model that will be used broadly across all of our customers, we do not build or train these models in such a way that they could learn, memorise, or be able to reproduce some part of Customer Data This feels so full of subtle qualifiers and weasel words that it generates far more distrust than trust. It only refers to models used "broadly across all" customers - so if it's (a) not used "broadly" or (b) only…

> They need to reword this. Whoever wrote it is a liability Sounds like it’s been written specifically to avoid liability.

I'm sure it was lawyers. It's always lawyers.

Re: Slack AI Training with Customer Data

#103
post #61

> We offer Customers a choice around these practices. If you want to exclude your Customer Data from helping train Slack global models, you can opt out. If you opt out, Customer Data on your workspace will only be used to improve the experience on your own workspace and you will still enjoy all of the benefits of our globally trained AI/ML models without contributing to the underlying models. Why would anyone not opt…

I’d be surprised if more 1% opt out.

Re: Slack AI Training with Customer Data

#106

The incentive for first party tool providers to do this is going to be huge, whether its Slack, Google, Microsoft, or really any other SaaS tool. Ultimately, if business want to avoid getting commoditized by their vendors, they need be in control of their data, and their AI strategy. And that probably ultimately means turning off all of these small-utility-very-expensive-and-might-ruin-your-business features, and act…

"commoditized by their vendors" is exactly the phrase I was looking for. It's why I wanted my co to self-host Mattermost instead of using Slack.

Re: Slack AI Training with Customer Data

#107

Earlier quoted context omitted.

Why shouldn’t AI be able to shitpost too? At the very least, and much more importantly, AI should be able to recognize shitposting.

This is the crux of it, and where I'm wondering if I'm missing something. Can it, today? My understanding is it cannot discern reality from fiction, thus "hallucinations" (a misnomer because it implies awareness, which these probability models lack).

The poorly named hallucinations are creation of ideas from provided prompts, which ideas are not grounded in reality. It isn't the mistaken adjudication of the reality of a provided prompt.

Re: Slack AI Training with Customer Data

#108

Story time. I was at a VC conference last year and if I learned nothing else there, I learned how to spell "AI". Every single exhibitor just about had their signage proudly proclaiming their capabilities in this area, but one in particular struck me. They were touting the API integrations they could offer to train their "Enterprise AI"/LLM, and among those integrations were things like M365, Slack, etc. It struck me…

The sheer scale of data on the long tail. Sure, the head is already a trash pile and has been for decades now, but there is plenty of non-monetized information all over the internet that is barely linked to or otherwise discoverable.

It does not matter how hard they try, nothing will rival the CommonCrawl treasure trove except maybe Google's index itself.

Re: Slack AI Training with Customer Data

#109

>Our mission is to build a product that makes work life simpler, more pleasant and more productive. I know it would be impossible but I wish we go back to the days when we didn't have Slack (or tools alike). Our Slack is a cesspool of people complaining, talking behind other people's backs, echo chamber of negativity etc. That probably speaks more to the overall culture of the company, but Slack certainly doesn't hel…

Your company sucks. I’ve used slack at four workplaces and it’s not been at all like that. A previous company had mailing lists and they were toxic as you describe. The tool was not the issue.

Re: Slack AI Training with Customer Data

#110
post #79

Earlier quoted context omitted.

This assumes everything in the training data set is accurate. Sometimes people are wrong, obtuse, sarcastic, etc. LLM's don't have any way of detecting or accounting for this, do they? That output, then being used to train other LLM's, just creates an ouroboros of AI generated dogshit.

LLMs are state-of-the-art at detecting sarcasm. It won't help if the data is just wrong though. Edit: https://arxiv.org/abs/2312.03706 Human performance on this benchmark (detecting sarcasm in Reddit comments) was 0.82, a BERT-based LLM scored 0.79. https://arxiv.org/abd/2106.05752 LSTM, 98% at detecting sarcasm in a Project Gutenburg-based dataset.

I literally can't tell if you're being sarcastic or not.
Post reply on HN