Live data from Hacker News

If you have a Claude account, they're going to train on your data moving forward

old.reddit.com

161–170 of 235 posts

Re: If you have a Claude account, they're going to train on your data moving forward

#161

To be honest, these companies already stole terabytes of data and don't even disclose their dataset, so you have to assume they'll steal and train at anything you throw at them

"Reading stuff freely posted on the internet" constitutes stealing now? Seems like an excessively draconian interpretation of property rights.

Forgot the 82TB of torrented books Meta has been using for training? I mean, yeah, it’s Meta. No surprise. But I won’t believe for one second that the other players didn’t do a similar thing. They just haven’t been caught yet.

Re: If you have a Claude account, they're going to train on your data moving forward

#162

I’m fine with that.

you are fine with paying, 20, 90 or 200 euros a month AND having your data mined? i must be getting old...

You have it backwards. I'm paying $200/month and so I want the thing to reflect more what I want than what the general public wants. They'd better be mining my data, despite that being a term I haven't heard in over a decade.

Some people were upset that Google Maps would just take the data that contributors give it for free. My problem was different. I use Google Maps and I want a way to correct it. I don't want to be paid for this. I want the tool I'm using to be correctable by me. The more I pay for it, the more I want it to be editable by me. I don't want compensation. I want it to be better. And I can make it better.

It's sort of why we picked Kong at a different company. Open source core meant that we could edit stuff we didn't like. In fact, considering that we paid, we wanted them to upstream what we changed.

Re: If you have a Claude account, they're going to train on your data moving forward

#163

Earlier quoted context omitted.

Use gpt4all or one of the other locally-hosted AI chatbots. Download a model and see if it works for you. It won't be as good as the latest models out there presumably, but at least you're not sending any chat data anywhere.

I'm really disheartened by all this shrugging of shoulders for a company mining user data. On hackernews... Anthropic initally positioned itself as a company with privacy and data protection as a priority, and had a real chance to claim moral high ground compared to its competitors.

I'm not shrugging my shoulders about this. I already don't use online LLMs for anything important because of privacy concerns already (i.e. I assume anything that lands on their server is fair game for them to do what they like with it). If you do have privacy concerns about your data, host it locally. I have played around with self-hosted LLMs and they work for what I need, which is mostly playing around with them for entertainment.

I'm careful about what data of mine lands on someone else's server, so I'm not a fan of this even without the dark patterns.

Re: If you have a Claude account, they're going to train on your data moving forward

#164

To be honest, these companies already stole terabytes of data and don't even disclose their dataset, so you have to assume they'll steal and train at anything you throw at them

"Reading stuff freely posted on the internet" constitutes stealing now? Seems like an excessively draconian interpretation of property rights.

[deleted]

Re: If you have a Claude account, they're going to train on your data moving forward

#165

To be honest, these companies already stole terabytes of data and don't even disclose their dataset, so you have to assume they'll steal and train at anything you throw at them

"Reading stuff freely posted on the internet" constitutes stealing now? Seems like an excessively draconian interpretation of property rights.

This is a quintessential bad faith comment.

The reference to terabytes of stolen data refers to copyrighted material. I think you know this but chose to frame it as "stuff freely posted on the internet" in order to mislead and strawman the other comment.

Re: If you have a Claude account, they're going to train on your data moving forward

#166

Earlier quoted context omitted.

"Reading stuff freely posted on the internet" constitutes stealing now? Seems like an excessively draconian interpretation of property rights.

This is a quintessential bad faith comment. The reference to terabytes of stolen data refers to copyrighted material. I think you know this but chose to frame it as "stuff freely posted on the internet" in order to mislead and strawman the other comment.

I meant it exactly as I said it. I do not agree that any theft occurred, either in law or in spirit, and I believe that reinterpretation of intellectual-property law in order to make it a crime would cause significant harm, greatly outweighing the benefits, as has been the case with every other expansion of intellectual property law I have seen.

Re: If you have a Claude account, they're going to train on your data moving forward

#167

Earlier quoted context omitted.

"Reading stuff freely posted on the internet" is also very different from a business having machines consume large volumes of data posted on the Internet for the purpose of generating value for them without compensating the creators. I'm not making a value judgement one way or the other, but "reading stuff freely posted on the Internet" is an oversimplification.

Okay, but "stealing" is also an oversimplification, to the point of absurdity. It makes no sense to put stuff up on the internet where it can freely be downloaded by anyone at any time, by people who are then free to do whatever they like with it on their own hardware, then complain that people have downloaded that stuff and done what they liked with it on their own hardware. "Having machines consume large volumes of…

That is not at all how the internet works. Try to download music from Napster and Lars will sue your ass.

Re: If you have a Claude account, they're going to train on your data moving forward

#168
post #33

So, I guess they run out of data to train on ... I wonder on how much they can rely on the data and what kind of "knowledge" they can extract. I never give feedback and most time (let's say 5 out of 6) the result cc produce it simply wrong. How can they know the result is valuable or not?

How can they know anything they train on is valuable? At the end of the day it doesn't matter. You got the wrong answer and didn't complain, so why would they care?

In general, I think any human generated content in pre-2022 is valuable because someone did some kind of validation (think about stack overflow answer with user confirming that a specific answer fixed their problem).

If they start to feed the next model with LLM generated crap, the overall performance will drop and instead of getting a useful answer 1 of 5 it will be 1 of 10(?) and probably a lot of us will cancel the subscription ... so in the end I think it matters.

Re: If you have a Claude account, they're going to train on your data moving forward

#169
post #133
post #85

Earlier quoted context omitted.

>Why would you assume anything? Because they already used data without permission on a much larger scale, so it's a perfectly logical assumption that they would continue doing so with their users?

I don't think that logically makes sense. Training on everything you can publicly scrape from the internet is a very different thing from training on data that your users submit directly to your service.

>Training on everything you can publicly scrape from the internet is a very different thing from training on data that your users submit directly to your service.

Yes. It's way easier and cheaper when the data comes to you instead of having to scrape everything elsewhere.

Re: If you have a Claude account, they're going to train on your data moving forward

#170

Earlier quoted context omitted.

This is a quintessential bad faith comment. The reference to terabytes of stolen data refers to copyrighted material. I think you know this but chose to frame it as "stuff freely posted on the internet" in order to mislead and strawman the other comment.

I meant it exactly as I said it. I do not agree that any theft occurred, either in law or in spirit, and I believe that reinterpretation of intellectual-property law in order to make it a crime would cause significant harm, greatly outweighing the benefits, as has been the case with every other expansion of intellectual property law I have seen.

Anthropic downloaded books from Library Genesis and The Pirate Library mirror. This is factual and reported on from court documents.

What’s the angle that describes this as fair use?

[0] https://www.businessinsider.com/anthropic-cut-pirated-millio...

Post reply on HN