Live data from Hacker News

Microsoft, OpenAI sued for ChatGPT 'privacy violations'

theregister.com

181–190 of 231 posts

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#181

Earlier quoted context omitted.

1. How would this not make tools like Github Copilot exorbitantly expensive? Why should I have to pay a tax to everyone else in the United States to use something that was disproportionately trained on my own data? 2. Given that the internet is global, is every country supposed to make their own versions of this? Will I have to pay the EU tax to use models that might have been trained on data that Europeans posted on…

To your first question, it would incentivize training of models on one's own data exclusively -- companies could train something like Copilot on their own code, for instance. To your second question, there's no way to have an international policy like this so yes each jurisdiction would do it independently -- just as they do with thousands of other similar things.

Also regarding international policy - good luck getting Chinese citizens to pay the US AI tax. Effectively you'd be nerfing anyone under US jurisdiction

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#182

Earlier quoted context omitted.

Anyone can read your blog and then post their own blog post using knowledge they learned while reading yours. ChatGPT "learned" from your blog that same way

Anthropomorphizing that it "learned" is disingenuous and I expect better from the HN crowd. If ChatGPT regurgitates verbatim or nearly verbatim, something it slurped up from OP's blog, is that not plagiarism? Where do you draw the line? Where would a reasonable person draw the line?

But ChatGPT doesn’t spit out verbatim from the blog.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#183
post #35

>For the 16 plaintiffs, the complaint indicates that they used ChatGPT, as well as other internet services like Reddit, and expected that their digital interactions would not be incorporated into an AI model. I don't expect this lawsuit to lead anywhere. But if it does, I hope it leads to some clear laws regarding data privacy and how TOS is binding. The recent ruling regarding web scraping makes the case against Ope…

Regardless of access rights to the data, I've yet to read a compelling argument why LLMs are even derivative works. You can't identify your Reddit comment in a ChatGPT conversation. How is it any different than a human learning English by reading Reddit? That human wouldn't be violating copyright every time they said a phrase that was repeated by hundreds of Redditors. My favorite LLM analogy so far is the "lossy jpe…

I've been thinking of the output as fanfiction/fan art. It shares many of the same complications regarding the ownership of ideas, commerical intent of writing, competition, and copyright. Fanfiction is generally a protected form of expression, but requires the work to be "transformative". Unlike with parodies and critisisms, fanfiction can be much harder to distinguish from original work. From that perspective, a large amount of the output of LLMs is so generic, that it's not possible to attribute it to one person. It's like trying to find the original author of "Once upon a time".

https://theinnisherald.com/the-other-once-upon-a-times-a-his...

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#184
post #86
post #77

Earlier quoted context omitted.

>You don’t get to make information publicly available. But not publicly available. But we do? Open sourcing something with caveats is common. This code is public BUT not for commercial use. This code is public BUT you must display attribution etc. Sure, blogposts are unlicensed (that I know) but the idea of something publicly available being held to restrictions is nothing new.

Do you allow commercial employees to read the code and incorporate knowledge obtained from the code into their brains?

Why is it that people keep on flogging dead horses?

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#185

Earlier quoted context omitted.

To your first question, it would incentivize training of models on one's own data exclusively -- companies could train something like Copilot on their own code, for instance. To your second question, there's no way to have an international policy like this so yes each jurisdiction would do it independently -- just as they do with thousands of other similar things.

I don't think a model trained on a single company's data would be nearly as helpful as a model trained on all publicly licensed code on the internet. But suppose it were... What if I'm not a massive corporation with millions of lines of code to train on and I want to pay for an AI coding assistant? Doesn't this make it effectively illegal for me to purchase such a product for a reasonable price when big companies wil…

The policy would exempt all except big companies from the fees. So if you set up your own, you don't pay. And the effect of the revenue threshold creating an advantage for small businesses is commonplace in policy across the board -- SMBs don't have many of the same costs and obligations as larger companies.

And this would not prevent you from explicitly licensing your code or writing to let people to train on it. But what it would do is say that if someone didn't explicitly license it then it is covered under the policy.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#186
post #98

Earlier quoted context omitted.

There is a difference between learning from your work and copying your work. You are entitled to control it's distribution and use. You are not entitled to control it's influence and effects.

I think you've made up an irrelevant argument. The work has been incorporated into a commercial product, intentionally, under the control of someone else. Software isn't humans that pay taxes, appear in court, have rights, etc.

No, the work has not been. The impression that the work leaves on a neural network has been though.

AIs are not massive repositories of harvested data. The models are relatively small (<20GB).

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#187

Earlier quoted context omitted.

To your first question, it would incentivize training of models on one's own data exclusively -- companies could train something like Copilot on their own code, for instance. To your second question, there's no way to have an international policy like this so yes each jurisdiction would do it independently -- just as they do with thousands of other similar things.

Also regarding international policy - good luck getting Chinese citizens to pay the US AI tax. Effectively you'd be nerfing anyone under US jurisdiction

Not really -- it's the same as selling any service into the US. Yes people cheat on, say, sales tax, just like Amazon did in the early years, but eventually once big enough companies end up having to adhere to the policy.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#188
post #172

Earlier quoted context omitted.

Anyone who reads your blog is "ingesting content" from it. That is presumably the purpose of your blog in the first place. Whether that content is used to train a human mind or an artificial one is probably not up to you as the author.

This type of comments can be seen every single time a thread about LLM, or OpenAI or some such comes up. And it adds nothing. I'm sorry but saying "Whether that content is used to train a human mind or an artificial one is probably not up to you " may be worse than saying nothing at all. First because it shows enough doubt on whether it's up to the authors of content (IP laws, fair use, intent of the use, and many th…

I disagree, and though the GP maybe didn't have this sentiment, my personal view is that intellectual property is a bunch of crap and just because there are laws around it in our capitalist society doesn't mean that the laws are moral/just/ethical/good. IP is constantly ingested and transformed which is exactly what LLMs are doing. The fact that ChatGPT can't even accurately reproduce data from its training (it gets basic facts/dates/quotes wrong all the time) really reinforces that it's not infringing on anyone's IP.

If you're tired of responding to these comments then stop. It's the internet, everyone is at different places in exploring topics and having discussions. Don't poo-poo on someone else's journey and instead move on with your day. There is no required reading (other than TFA) on hacker news.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#189

Earlier quoted context omitted.

> I hope it leads to some clear laws regarding data privacy and how TOS is binding I hope it leads to more people realizing that a TOS doesnt override their individual rights and that the legal system works to support them.

One individual right is the right to sign away other rights in exchange for products and services.

There are limits to that -- to signing away rights. In the US You can't sign yourself into slavery. You can't sell the right to have someone kill you.

There's sort of an exception for military service, but even soldiers have acess to military courts.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#190

Earlier quoted context omitted.

There is a difference between learning from your work and copying your work. You are entitled to control it's distribution and use. You are not entitled to control it's influence and effects.

Actually, people have been successfully sued for plagiarizing other works because they had internalized it and accidentally regurgitated it. So. The fact that content runs through a human brain doesn’t necessarily cleanse it from copyright concerns.

There is no "actually" because you are still addressing distribution. It wouldn't be hard to have another AI that analyzes outputs for copywriter infringement and culls them as necessary.

Would that satisfy you?

Post reply on HN