Live data from Hacker News

Can I opt out of my input or output data being used for training?

help.mistral.ai

251–260 of 262 posts

Re: Can I opt out of my input or output data being used for training?

#251
post #222

Earlier quoted context omitted.

Have you tried Lumo by Proton? It is E2EE and has projects. A bit more expensive than duck.ai but I found the model to be good enough

How could a cloud LLM be E2EE?

Google has a paper on this: https://arxiv.org/html/2409.19134v5

I believe Apple has something similar too.

Re: Can I opt out of my input or output data being used for training?

#252
post #2

Context: After careful research our organization preferred a European partner with good central privacy controls. We landed on Mistral, after being disappointed that the Pro tier was opt-in to training on prompts by default we switched up to the Team tier which provides an organization dashboard with some relevant settings. As we did that Mistral changed these options and the Team tier was now also opt-in by default…

Buried by the distraction of the semantics of "opt in" vs "opt out", the more interesting conflict isn't addressed in the sibling threads.

> and at the same time seemed to have lost the ability to centrally disable training on prompts for your entire organization

vs

> Vibe (Teams): Administrators can disable data training usage for the entire organization.

Is the document out of date? Or did Mistral reinstate this ability after the fact? What's the story here? Others seem to be stating they have and have had the ability to opt out of training centrally for a long time.

EDIT: Oh, it is addressed just a bit hard to find with all the opt in/out explanations: https://news.ycombinator.com/item?id=49549102

Re: Can I opt out of my input or output data being used for training?

#253
Can anyone shed light on why there's a button for copying specifically for LLMs ("Copy for LLM") and why there's also an "Open in Claude" option on these articles?

Regarding the former, there's no normal Copy button so presumably this just copies the content of the article. Not sure why it needs to be clarified that its for LLMs.

Re: Can I opt out of my input or output data being used for training?

#254
post #140

You have to be rather naive if you don't think these companies don't simply train on your prompts with or without your consent. They literally scrape everything - legal or not - and claim its fair use to train on, including straight piracy The idea that they'll steal from everyone except you is just wishful thinking

I think innocent-until-proven-guilty is the correct approach in general but... It's also immature to assume your data will stay private in the long term. These AI companies are immensely valuable targets. When they get breached, their datasets will inevitably penetrate public datasets. And why would any AI company refuse to train on 'public' data?

innocent until proven guilty is definitely wrong approach vs tech giants who time and time again prove they will do anything unethical if only they can theoretize how to get away with it or the fine is low enough

Re: Can I opt out of my input or output data being used for training?

#255

Earlier quoted context omitted.

I don’t care about what they do, I only care about what they say, that way, when compliance requires us to use LLMs that don’t train on inputs, I can point to that policy and continue on using the LLM. If they are secretly scraping input prompts, that no longer becomes my problem. Someday, a massive lawsuit comes down, and everyone gets to play the victim. You’re naive for thinking we don’t know how this really works…

What they say is not the same as what they wrote in the ToS you didn't read. If they stole your stuff and used in to train their model, what does it help if there is a lawsuit some years later which you most likely have no benefit from?

> If they stole your stuff and used in to train their model, what does it help if there is a lawsuit some years later which you most likely have no benefit from?

The point is that if someone accuses you of leaking data you can sue them for not abiding by contract, not that the data is not used for training

Re: Can I opt out of my input or output data being used for training?

#256
post #46

Earlier quoted context omitted.

I don't have that toggle (but could indeed have sworn I saw it earlier). Sorry, I have always thought that "opting in" is, "opting for the presented option" and opting out is "opting out of it", so opting out [of sharing prompts for training] is choosing to not share, but apparently I was wrong my whole life. I'm not a native speaker, and I think most people here (in my country) would interpret this the way I do? Wei…

> I don't have that toggle (but could indeed have sworn I saw it earlier). You don't see "Allow the use of your interactions with Vibe to train Mistral’s AI models" at https://admin.mistral.ai/vibe/privacy ? And "Allow the use of your API calls to train Mistral’s AI models" at https://admin.mistral.ai/plateforme/privacy ?

Ok it (org wide toggle) just (24h after posting this) returned and the docs were changed again! They now include the text:

"Vibe (Teams): Administrators can disable data training usage for the entire organization."

That’s nice! Except that it was on for my entire org so I’ll be checking if they didn’t store anything first thing tomorrow. I wonder what changed their minds.

Re: Can I opt out of my input or output data being used for training?

#257
post #2

Context: After careful research our organization preferred a European partner with good central privacy controls. We landed on Mistral, after being disappointed that the Pro tier was opt-in to training on prompts by default we switched up to the Team tier which provides an organization dashboard with some relevant settings. As we did that Mistral changed these options and the Team tier was now also opt-in by default…

Buried by the distraction of the semantics of "opt in" vs "opt out", the more interesting conflict isn't addressed in the sibling threads. > and at the same time seemed to have lost the ability to centrally disable training on prompts for your entire organization vs > Vibe (Teams): Administrators can disable data training usage for the entire organization. Is the document out of date? Or did Mistral reinstate this ab…

You are right! They must have just now switched this back, I also have the toggle now! It was turned on sadly but it’s something.

Re: Can I opt out of my input or output data being used for training?

#258
post #2

Context: After careful research our organization preferred a European partner with good central privacy controls. We landed on Mistral, after being disappointed that the Pro tier was opt-in to training on prompts by default we switched up to the Team tier which provides an organization dashboard with some relevant settings. As we did that Mistral changed these options and the Team tier was now also opt-in by default…

Just now the sentence:

“Vibe (Teams): Administrators can disable data training usage for the entire organization.”

Was added to tfa. And I now see an org wide toggle where there was none before! Sadly it was on so I hope no users submitted stuff in the mean time, but it’s something!

Re: Can I opt out of my input or output data being used for training?

#259
post #46

Earlier quoted context omitted.

I don't have that toggle (but could indeed have sworn I saw it earlier). Sorry, I have always thought that "opting in" is, "opting for the presented option" and opting out is "opting out of it", so opting out [of sharing prompts for training] is choosing to not share, but apparently I was wrong my whole life. I'm not a native speaker, and I think most people here (in my country) would interpret this the way I do? Wei…

> I don't have that toggle (but could indeed have sworn I saw it earlier). You don't see "Allow the use of your interactions with Vibe to train Mistral’s AI models" at https://admin.mistral.ai/vibe/privacy ? And "Allow the use of your API calls to train Mistral’s AI models" at https://admin.mistral.ai/plateforme/privacy ?

[deleted]

Re: Can I opt out of my input or output data being used for training?

#260

Earlier quoted context omitted.

that was accidental? this would be straight up fraud.

You can structure it so that it becomes accidental. 1. Ensure security barriers are weak or honor based. 2. Put individual researchers under a lot of pressure. 3. If you get caught, blame the weak barriers, or the individual researcher. Basically setup the incentive structure to incentivize researchers sticking their mittens in the private cookie jar while putting the cookie jar in a dark unmonitored/unsecured room w…

Yeah, the steps follow exactly what happened at VW with DieselGate. The diesel emissions lies were found out because some enterprising person set up an emissions testing system and drove the car in real world scenarios with it to verify the claimed emissions.

There's no reliable way to verify a foundation model has been trained on a particular piece of proprietary data. If an API key is ingested, hopefully the foundation model is wrapped in enough moderation that the raw API key oberserved during training is not recited verbatim in the output.

Post reply on HN