Live data from Hacker News

Microsoft, OpenAI sued for ChatGPT 'privacy violations'

theregister.com

211–220 of 231 posts

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#211

Earlier quoted context omitted.

If you go read a book, memorize it, write it down later in a substantively similar form, and share it freely or sell it — yes, you might get into copyright trouble. It has happened before and it is at best a tricky gray area. If you pick up a book and learn a fact, then yeah, you’re allowed to share that fact. It’s weird that this topic keeps devolving into a form of “so what, it’s illegal for me to learn things?” Be…

> You have a different set of rights than ChatGPT. Gods, no. Where did you get that from?

Are you a human being? A citizen of some country? If so you definitely have a different set of rights than ChatGPT.

Those might not be a problem regarding this specific case, but the case can easily be made that it ought to be.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#212

Earlier quoted context omitted.

Anyone can read your blog and then post their own blog post using knowledge they learned while reading yours. ChatGPT "learned" from your blog that same way

Anthropomorphizing that it "learned" is disingenuous and I expect better from the HN crowd. If ChatGPT regurgitates verbatim or nearly verbatim, something it slurped up from OP's blog, is that not plagiarism? Where do you draw the line? Where would a reasonable person draw the line?

Actually I fear that people that say this are doing worse than anthropomorphizing.

Often rather than claiming human aspects to the machine, they are going further, and claiming machine aspects to the human.

Using mechanistic analogies for explaining the human body or mind isn't new, but as machines become better and better at imitating humans, those analogies become more seductive.

That's my rant; the danger with 'AI' isn't so much that humans are enslaved by machines, but that we enslave each other -- or dehumanize each other -- with machines.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#213
post #44

I mean, it ingested all of the content from my blog. Without my permission. It's not a major part of their corpus of data, but still -- I wasn't asked and I don't really care to donate work to large corporations like that. So the technology is cool, but I'm firmly of the stance that they cut corners and trampled peoples' rights to get a product out the door. I wouldn't be entirely unhappy if this iteration of these p…

If you put your content on a billboard, what expectation should you have that you can control who reads it?

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#214
post #172

Earlier quoted context omitted.

Anyone who reads your blog is "ingesting content" from it. That is presumably the purpose of your blog in the first place. Whether that content is used to train a human mind or an artificial one is probably not up to you as the author.

This type of comments can be seen every single time a thread about LLM, or OpenAI or some such comes up. And it adds nothing. I'm sorry but saying "Whether that content is used to train a human mind or an artificial one is probably not up to you " may be worse than saying nothing at all. First because it shows enough doubt on whether it's up to the authors of content (IP laws, fair use, intent of the use, and many th…

I think you’ve missed the point. Copyright laws prevent others from copying your work without permission. (Hence the name.) Copyright laws say nothing about who can read your work.

If you want to prevent a web spider from scraping your blog, use a captcha or robots.txt. Copyright law doesn’t apply to this scenario.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#215

Earlier quoted context omitted.

I can personally memorize and recite copyrighted works all I want, but when ChatGPT does it then it’s in a commercial context and they’re liable to be sued for infringement. If you ask ChatGPT the rules for D&D, the private sourcebooks are all in there.

> and recite copyrighted works all I want ...wait, isn't that false? legitimately asking. or is it because it was done by a corporation that makes it illegal? im thinking of how restaurants dont sing happy birthday and fair use restrictions etc

Yes, I was overly broad and there are restrictions on saying/copying memorized material.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#216
post #64

Earlier quoted context omitted.

I'd say this is more like if someone automated taking pictures of every flyer and missing pet poster people put up on a lightpole and saved it to a database. There's more deliberate action when you post something on a public online form than just existing in a place outside of your house. Especially considering you've always had the option to use reddit anonymously anyway.

>use reddit anon.... Read, yes - post no. And - you can no longer create an account that is not tied to an email...

OpenAI didn't have access to every poster's email when they crawled reddit. If you're making posts or have an account name that are easily tied back to your personal identity, that's on you. But you could make an account with any random username you wanted, that keeps you anonymous as far as OpenAI is concerned.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#217
post #198

Earlier quoted context omitted.

No, the work has not been. The impression that the work leaves on a neural network has been though. AIs are not massive repositories of harvested data. The models are relatively small (<20GB).

A resized, smaller, or encoded version of an image is still subject to copyright. Calling an encoding an 'impression' is deceitful.

It's none of the those things, these models train on petabytes of data. They store relationships of objects to each other, not objects themselves.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#218
post #202
post #198

Earlier quoted context omitted.

A resized, smaller, or encoded version of an image is still subject to copyright. Calling an encoding an 'impression' is deceitful.

Not always. https://www.pinsentmasons.com/out-law/news/google-thumbnails... > A US court ruled this week that Google's creation and display of thumbnail images does not infringe copyright. It also said that Google was not responsible for the copyright violations of other sites which it frames and links to.

Part of this ruling is about how the images are used -- Fair use -- not just that they were stored in a particular way. If Google was using the smaller versions of the images (thumbnails) in other ways, it could have been infringing.

> The Court said that Google did claim fair use, and that whether or not use was fair depended on four factors: the purpose and character of the use, including whether such use is of a commercial nature or is for non-profit educational purposes; the nature of the copyrighted work; the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and the effect of the use upon the potential market for or value of the copyrighted work.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#219
post #198

Earlier quoted context omitted.

A resized, smaller, or encoded version of an image is still subject to copyright. Calling an encoding an 'impression' is deceitful.

It's none of the those things, these models train on petabytes of data. They store relationships of objects to each other, not objects themselves.

[deleted]

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#220

Earlier quoted context omitted.

> People put stuff up on the Internet with the expectation of its consumption by human minds Then people obviously aren’t aware that bots have been indexing web pages and showing summarized information without going to the web page for three decades.

I think it's a bit intellectually dishonest to claim an equivalence between content indexing for search engines and machine learning for LLMs. They might share an underlying harvesting technique, but their uses -- indexing for information accessibility vs automatic content production are qualitatively different. Further, almost every site has had an e.g. robots.txt which has permitted content harvesting only for cert…

How is it not content production when I search for something on Google and get a box with similar questions and summarizes the answer.

So you’re okay with Google making money off of your content. But not OpenAI?

Post reply on HN