Live data from Hacker News

More writers sue OpenAI for copyright infringement over AI training

reuters.com

21–30 of 68 posts

Re: More writers sue OpenAI for copyright infringement over AI training

#21

Earlier quoted context omitted.

> Assuming they used that data... That's the key part. You haven't yet proved they have actually used your content for anything (other than, potentially, read the license to decide if they should include or discard from their training set). But in practice we'll never know for sure if they are respecting the terms of licenses until 1) this is tested in court, or 2) there's some internal leak that points into either d…

I expect that OpenAI would concede that they used the data in any court case immediately to get that issue off the table, I really don't think they have a strong interest in foot-dragging on this stuff, right? I would think OpenAI wants the thornier legal issues actually settled so that the whole ecosystem can grow within those terms & they can lobby for the legal changes they need/want?

The alternative would be discovery on that issue, which they may want to avoid.

Re: More writers sue OpenAI for copyright infringement over AI training

#22
All these lawsuits will die. Why?

Because people train on corpuses of data all the time, without a license or any attribution.

Every piece of text a writer reads is training that writer. Every image an artist sees helps to train that artist. Every sound a musician hears is training that musician.

That doesn't mean they can't exclude their works from training via a license going foreward. But that becomes an enforcement problem.

Re: More writers sue OpenAI for copyright infringement over AI training

#23
post #5

Earlier quoted context omitted.

Not at all like piracy. When someone pirates a book, they're replacing the original without consent or remuneration to the copyright holders. When you train an AI on the contents of a book, you're not replacing it. If someone is interested in the content, they still need to buy it. Using ChatGPT is not a substitute. If it is, they're gonna have to prove it in court, but I doubt they'll be able to.

If you can ask ChatGPT about any book contents, you don't need to get the book, and if you don't need to get the book then author got robbed, ClosedAI/MS profited.

I can also go to the Wikipedia article about any book of note and get a plot or other summary and other information about the book. If that’s the reason for buying the book, “the author got robbed.”

Re: More writers sue OpenAI for copyright infringement over AI training

#24
post #22

All these lawsuits will die. Why? Because people train on corpuses of data all the time, without a license or any attribution. Every piece of text a writer reads is training that writer. Every image an artist sees helps to train that artist. Every sound a musician hears is training that musician. That doesn't mean they can't exclude their works from training via a license going foreward. But that becomes an enforceme…

IIRC courts have already ruled AI-generated works cannot have copyright. So there is already a legal distinction between a human and a model creating works.

I also doubt "humans are just a larger Markov chain than the LLM and they're allowed to" will hold up in court.

Re: More writers sue OpenAI for copyright infringement over AI training

#25
post #22

All these lawsuits will die. Why? Because people train on corpuses of data all the time, without a license or any attribution. Every piece of text a writer reads is training that writer. Every image an artist sees helps to train that artist. Every sound a musician hears is training that musician. That doesn't mean they can't exclude their works from training via a license going foreward. But that becomes an enforceme…

People can do a lot of things that we don't legally allow machines or automation to do.

Re: More writers sue OpenAI for copyright infringement over AI training

#26
post #22

All these lawsuits will die. Why? Because people train on corpuses of data all the time, without a license or any attribution. Every piece of text a writer reads is training that writer. Every image an artist sees helps to train that artist. Every sound a musician hears is training that musician. That doesn't mean they can't exclude their works from training via a license going foreward. But that becomes an enforceme…

People can do a lot of things that we don't legally allow machines or automation to do.

Like what? The only things I can think of relate to quality/safety (eg. drivers or lawyers).

Re: More writers sue OpenAI for copyright infringement over AI training

#27
post #22

All these lawsuits will die. Why? Because people train on corpuses of data all the time, without a license or any attribution. Every piece of text a writer reads is training that writer. Every image an artist sees helps to train that artist. Every sound a musician hears is training that musician. That doesn't mean they can't exclude their works from training via a license going foreward. But that becomes an enforceme…

IIRC courts have already ruled AI-generated works cannot have copyright. So there is already a legal distinction between a human and a model creating works. I also doubt "humans are just a larger Markov chain than the LLM and they're allowed to" will hold up in court.

I don’t see what eligibility to have works protected has to do with legality of learning.

I really hope “copyright can be used to prohibit reading and learning” does not hold up in court.

Copyright is, and should be, a protection from unauthorized reproduction. Extending it to protect the abstract ideas would be a disaster. And extending it to control stylistic learning would be even worse.

Re: More writers sue OpenAI for copyright infringement over AI training

#28

Random thought: my blog is licensed under a Creative Commons license [1] that allows you to use and transform my content as long as you give attribution and distribute your contributions under the same terms. I found the OpenAI bot scraping my blog recently. Assuming they used that data, when will they attribute me? [1] https://creativecommons.org/licenses/by-sa/4.0/

I believe that OpenAI is not required to attribute you if the output was produced by an OpenAI-operated AI model because the AI is not constrained by the Berne Convention treaty regime in the same way that people are.

I believe that this fact is and will be exploited to strip copyright and effectively transfer ownership using cleanroom/firewall techniques.

Re: More writers sue OpenAI for copyright infringement over AI training

#29

Random thought: my blog is licensed under a Creative Commons license [1] that allows you to use and transform my content as long as you give attribution and distribute your contributions under the same terms. I found the OpenAI bot scraping my blog recently. Assuming they used that data, when will they attribute me? [1] https://creativecommons.org/licenses/by-sa/4.0/

Well it all comes down to whether training an LLM is fair use or not. I think it is likely that courts rule it is transformative enough that training is allowed regardless of what terms you have for the use of content.

Re: More writers sue OpenAI for copyright infringement over AI training

#30
post #22

All these lawsuits will die. Why? Because people train on corpuses of data all the time, without a license or any attribution. Every piece of text a writer reads is training that writer. Every image an artist sees helps to train that artist. Every sound a musician hears is training that musician. That doesn't mean they can't exclude their works from training via a license going foreward. But that becomes an enforceme…

It's a good thing humans and computers are 2 wholly separate categories of things that have 0 things related to them other than computers being anthropomorphized by AI sycophants!
Post reply on HN