Live data from Hacker News

If you have a Claude account, they're going to train on your data moving forward

old.reddit.com

121–130 of 235 posts

Re: If you have a Claude account, they're going to train on your data moving forward

#121
post #103

Earlier quoted context omitted.

"Reading stuff freely posted on the internet" is also very different from a business having machines consume large volumes of data posted on the Internet for the purpose of generating value for them without compensating the creators. I'm not making a value judgement one way or the other, but "reading stuff freely posted on the Internet" is an oversimplification.

We didn't seem to mind when Google was doing it back in 1999, or Lycos, Altavista, etc before them... why do we care about the LLM companies doing it now?

Because they have terms of service they have to adhere to. We need laws to be lawful.

Re: If you have a Claude account, they're going to train on your data moving forward

#122

Title is misleading, they're now opt-out rather than opt-in to your data being used for training. All you have to do is flip a single switch in the options to turn it off, I don't understand why everyone is treating this as being such a big deal. Edit: I just logged in to opt out, they presented me with the switch directly. It was two clicks.

[Edit: turns out I've got it wrong, and 5-year retention only said to apply to the data they're allowed to train on. This changes things for me.] Personally, I don't mind training, as long as I have a say on the matter - and they have a switch for this. Opt-out is not exactly cool, but I've got the popup in my face, almost a month before the changes, and that's respectful enough for me. This said, I've just canceled…

Data retention appears to be predicated on opting in to allowing training. If you don't opt in, they retain it for the same 30 days they were already retaining it for. https://www.anthropic.com/news/updates-to-our-consumer-terms

Re: If you have a Claude account, they're going to train on your data moving forward

#123
post #120

To be honest, these companies already stole terabytes of data and don't even disclose their dataset, so you have to assume they'll steal and train at anything you throw at them

No you don't. You don't have to assume people are going to be bad! We should not normalize it either.

You don't have to assume people are going to be bad, but it's reasonable and prudent to expect it from people who have already shown themselves to be so (in this context).

I trust people until they give me cause to do otherwise.

Re: If you have a Claude account, they're going to train on your data moving forward

#124
post #103

Earlier quoted context omitted.

We didn't seem to mind when Google was doing it back in 1999, or Lycos, Altavista, etc before them... why do we care about the LLM companies doing it now?

I find LLMs extremely useful but I think the difference is that they regurgitate the content (not verbatim) instead of a link to it. This is not unlike how a human might tell their friend about it.

> This is not unlike how a human might tell their friend about it.

Is there someone who has read the whole internet? Can we all be there friend?

The entire basis of fair use is scale matters.

Re: If you have a Claude account, they're going to train on your data moving forward

#125
post #104

Earlier quoted context omitted.

Okay, but "stealing" is also an oversimplification, to the point of absurdity. It makes no sense to put stuff up on the internet where it can freely be downloaded by anyone at any time, by people who are then free to do whatever they like with it on their own hardware, then complain that people have downloaded that stuff and done what they liked with it on their own hardware. "Having machines consume large volumes of…

They are not free to do whatever they like, there are tomes of laws across all countries governing what someone can and cannot do with your intellectual property. Just because we didn't have the foresight to add in a "if by chance in the future someone invents artificial intelligence, that's not fair use" is a shame, but doesn't make what these companies are doing ethical or morale. I don't disagree regarding Google,…

I did say "free to do whatever they like on their own hardware", because intellectual property laws generally govern the transfer of such property rather than the use.

After seeing the harm done by the expansion of patent law to cover software algorithms, and the relentless abuse done under the DMCA, I am reflexively skeptical of any effort to expand intellectual property concepts.

Re: If you have a Claude account, they're going to train on your data moving forward

#126
post #120

Earlier quoted context omitted.

No you don't. You don't have to assume people are going to be bad! We should not normalize it either.

You don't have to assume people are going to be bad, but it's reasonable and prudent to expect it from people who have already shown themselves to be so (in this context). I trust people until they give me cause to do otherwise.

Training on personal data people thought was going to remain private vs. stuff out in public view (copyright or not), are two different magnitudes of ethics breaches. Opt OUT instead of Opt IN for this is CRAZY in my opinion. I hope that the reddit post is WRONG on that detail but I seriously doubt it.

I asked Claude: "If a company has a privacy policy and says they will not train on your data and then decides to change the policy in order "to make the models better for everyone." What should the terms be?"

The model suggests in the first paragraph or so EXPLICIT OPT IN. Not Opt OUT

Re: If you have a Claude account, they're going to train on your data moving forward

#127

Earlier quoted context omitted.

[Edit: turns out I've got it wrong, and 5-year retention only said to apply to the data they're allowed to train on. This changes things for me.] Personally, I don't mind training, as long as I have a say on the matter - and they have a switch for this. Opt-out is not exactly cool, but I've got the popup in my face, almost a month before the changes, and that's respectful enough for me. This said, I've just canceled…

Data retention appears to be predicated on opting in to allowing training. If you don't opt in, they retain it for the same 30 days they were already retaining it for. https://www.anthropic.com/news/updates-to-our-consumer-terms

Oh! Thank you!

That popup was confusing as hell then, because I've read and understood it as two separate points: I've got it that they're making training opt-out, and that they're changing data retention to 5 years, independent of each other. I got upset over this, and haven't really researched into the nuances - and turns out I've got it all wrong.

Appreciate your comment, it's really helpful!

I hope they change the language to make it clear 5 years only applies to the chats they're allowed to train models on.

(Weirdly, I can't find the word "years" anywhere on their Privacy Policy, and the only instance on the Consumer Terms of Service pages is about being of legal age over 18 years old.)

Re: If you have a Claude account, they're going to train on your data moving forward

#128

Excellent. What were they waiting for up to now?? I thought they already trained on my data. I assume they train, even hope that they train, even when they say they don't. People that want to be data privacy maximalists - fine, don't use their data. But there are people out there (myself) that are on the opposite end of the spectrum, and we are mostly ignored by the companies. Companies just assume people only ever w…

The fact you are not aware of abuse, or abuse has not yet happened to you, does not mean it isn't a problem for you. > The defaults are always "deny everything". This is definitely not true for a massive amount of things, I'm unsure how you're even arriving at this conclusion.

Maybe in the US. In the UK, I have found obstacles to data sharing codified in the UK law frustrating. I'm reasonably sure some people will have died because of this, that would not have died otherwise. "Otherwise" case being - if they could communicate with the NHS, similarly (via email, whatsapp) to how they communicate in their private and professional lives.

Within the UK NHS and UK private hospital care, these are my personal experiences.

1) Can't email my GP to pass information back-and-forth. GP withholds their email contact, I can't email them e.g. pictures of scans, or lab work reports. In theory they should have those already on their side. In practice they rarely do. The exchange of information goes sms->web link->web form->submit - for one single turn. There will be multiple turns. Most people just give up.

2) MRI scan private hospital made me jump 10 hops before sending me link, so I can download my MRI scans videos and pictures. Most people would have given up. There were several forks in the process where in retrospect could have delayed data DL even more.

3) Blood tests scheduling can't tell me back that scheduled blood test for a date failed. Apparently it's between too much to impossible for them to have my email address on record, and email me back that the test was scheduled, or the scheduling failed. And that I should re-run the process.

4) I would like to volunteer my data to benefit R&D in the NHS. I'm a user of medicinal services. I'm cognisant that all those are helping, but the process of establishing them relied on people unknown to me sharing very sensitive personal information. If it wasn't for those unknown to me people, I would be way worse off. I'd like to do the same, and be able to tell UK NHS "here are, my lab works reports, 100 GB of my DNA paid for by myself, my medical histories - take them all in, use them as you please."

In all cases vague mutterings of "data protection... GDPR..." have been relayed back as "reasons". I take it's mostly B/S. Yes there are obstacles, but the staff could work around if they wanted to. However there is a kernel of truth - it's easier for them to not try to share, it's less work and less risk, so the laws are used as a cover leaf. (in the worst case - an alibi for laziness.)

Re: If you have a Claude account, they're going to train on your data moving forward

#130
post #120

Earlier quoted context omitted.

No you don't. You don't have to assume people are going to be bad! We should not normalize it either.

You don't have to assume people are going to be bad, but it's reasonable and prudent to expect it from people who have already shown themselves to be so (in this context). I trust people until they give me cause to do otherwise.

No, nbulka is correct. People should not shrug off and accept things that are wrong just because it's to be expected. It's one of the worst things you can do because as already pointed out, it just normalizes wrong.
Post reply on HN