Live data from Hacker News

More writers sue OpenAI for copyright infringement over AI training

reuters.com

11–20 of 68 posts

Re: More writers sue OpenAI for copyright infringement over AI training

#11
post #5

Earlier quoted context omitted.

Not at all like piracy. When someone pirates a book, they're replacing the original without consent or remuneration to the copyright holders. When you train an AI on the contents of a book, you're not replacing it. If someone is interested in the content, they still need to buy it. Using ChatGPT is not a substitute. If it is, they're gonna have to prove it in court, but I doubt they'll be able to.

If you can ask ChatGPT about any book contents, you don't need to get the book, and if you don't need to get the book then author got robbed, ClosedAI/MS profited.

The ability to ask people what is contained within a book isn't obviously copyright infringement.

Merely summarizing info and attributing it to the source is the basic element of learning, for both machines and human beings.

These suits are necessary becsuse it's not clear where the line is, and if ChatGPTs functions actually cross it.

What is clear is that OpenAI is doing its best to avoid infringing anyone's copyright even if it is trivial for them to do so. They have the training data so they can simply output it word for word bypass the LLM. They don't do that and further restrain their LLM from making too long recitations.

If you can trick / manipulate the LLM into giving you too much then I say that infringement is on you.

Re: More writers sue OpenAI for copyright infringement over AI training

#12

Random thought: my blog is licensed under a Creative Commons license [1] that allows you to use and transform my content as long as you give attribution and distribute your contributions under the same terms. I found the OpenAI bot scraping my blog recently. Assuming they used that data, when will they attribute me? [1] https://creativecommons.org/licenses/by-sa/4.0/

Interesting question, continuing on this, since they probably used GPL-3 code with the Affero clause, do they have to open source GPT? (The Affero clause is I believe the more directly applicable license thingy, though CC by-sa should also work.)

https://www.gnu.org/licenses/agpl-3.0.en.html

Re: More writers sue OpenAI for copyright infringement over AI training

#13

Random thought: my blog is licensed under a Creative Commons license [1] that allows you to use and transform my content as long as you give attribution and distribute your contributions under the same terms. I found the OpenAI bot scraping my blog recently. Assuming they used that data, when will they attribute me? [1] https://creativecommons.org/licenses/by-sa/4.0/

Where's your attributions for all the words written in your comment ;P you remixed the words and grammar patterns from other people's creative common's licenses of other people writings!

Note: i'm declaring my comment license as https://creativecommons.org/licenses/by-sa/4.0/

So if you remix or transform my comment by responding it, please attribute to me your response.

Re: More writers sue OpenAI for copyright infringement over AI training

#14
post #5

Earlier quoted context omitted.

Not at all like piracy. When someone pirates a book, they're replacing the original without consent or remuneration to the copyright holders. When you train an AI on the contents of a book, you're not replacing it. If someone is interested in the content, they still need to buy it. Using ChatGPT is not a substitute. If it is, they're gonna have to prove it in court, but I doubt they'll be able to.

If you can ask ChatGPT about any book contents, you don't need to get the book, and if you don't need to get the book then author got robbed, ClosedAI/MS profited.

What sort of questions? How would you know what to ask it, unless maybe you have another source for the book?

Curiously, when I ask GPT-4 about some well-known but under-copyright book, it says it can't answer because of the copyright. For well-known books out of copyright such as Alice in Wonderland, it can recite passages but tends to get lost and start reciting another section or book at some point. Would be real frustrating to use as a substitute.

Re: More writers sue OpenAI for copyright infringement over AI training

#15
this is a great lawsuit! if you read the complaint, they catch OpenAI dead-to-rights .. asking about plot details with names from the books, asking to write a paragraph in the style of the author in that book, and a diversity of authors that shows social awareness.. great support for this from California

Re: More writers sue OpenAI for copyright infringement over AI training

#17

I'm miffed. I tried a couple characters from my books, and zilch: ===== who is dan markunas ChatGPT I'm sorry, but I don't have any information on a person named Dan Markunas in my database .... who is janet saunders ChatGPT I'm sorry, but I don't have any specific information about a person named Janet Saunders in my database, ===========

Your book was published somewhere mid-2021, right?

Re: More writers sue OpenAI for copyright infringement over AI training

#18

Random thought: my blog is licensed under a Creative Commons license [1] that allows you to use and transform my content as long as you give attribution and distribute your contributions under the same terms. I found the OpenAI bot scraping my blog recently. Assuming they used that data, when will they attribute me? [1] https://creativecommons.org/licenses/by-sa/4.0/

Where's your attributions for all the words written in your comment ;P you remixed the words and grammar patterns from other people's creative common's licenses of other people writings! Note: i'm declaring my comment license as https://creativecommons.org/licenses/by-sa/4.0/ So if you remix or transform my comment by responding it, please attribute to me your response.

which works specially?

profit vs non profit also makes a difference

Re: More writers sue OpenAI for copyright infringement over AI training

#19
post #5

Earlier quoted context omitted.

Not at all like piracy. When someone pirates a book, they're replacing the original without consent or remuneration to the copyright holders. When you train an AI on the contents of a book, you're not replacing it. If someone is interested in the content, they still need to buy it. Using ChatGPT is not a substitute. If it is, they're gonna have to prove it in court, but I doubt they'll be able to.

If you can ask ChatGPT about any book contents, you don't need to get the book, and if you don't need to get the book then author got robbed, ClosedAI/MS profited.

If you ask a human questions about a book, thereby avoiding having to buy the book for class, did you rob the author?

Re: More writers sue OpenAI for copyright infringement over AI training

#20

Random thought: my blog is licensed under a Creative Commons license [1] that allows you to use and transform my content as long as you give attribution and distribute your contributions under the same terms. I found the OpenAI bot scraping my blog recently. Assuming they used that data, when will they attribute me? [1] https://creativecommons.org/licenses/by-sa/4.0/

> Assuming they used that data... That's the key part. You haven't yet proved they have actually used your content for anything (other than, potentially, read the license to decide if they should include or discard from their training set). But in practice we'll never know for sure if they are respecting the terms of licenses until 1) this is tested in court, or 2) there's some internal leak that points into either d…

I expect that OpenAI would concede that they used the data in any court case immediately to get that issue off the table, I really don't think they have a strong interest in foot-dragging on this stuff, right?

I would think OpenAI wants the thornier legal issues actually settled so that the whole ecosystem can grow within those terms & they can lobby for the legal changes they need/want?

Post reply on HN