Live data from Hacker News

More writers sue OpenAI for copyright infringement over AI training

reuters.com

41–50 of 68 posts

Re: More writers sue OpenAI for copyright infringement over AI training

#42

Earlier quoted context omitted.

Where's your attributions for all the words written in your comment ;P you remixed the words and grammar patterns from other people's creative common's licenses of other people writings! Note: i'm declaring my comment license as https://creativecommons.org/licenses/by-sa/4.0/ So if you remix or transform my comment by responding it, please attribute to me your response.

Humans are not LLMs trained and operated by a company for profit. Your argument is that LLMs hold all the same basic rights as humans but they hold (and should hold) exactly none.

That might be the implied argument, but the explicit argument appears to be the licensing a grouping of words as a work and then declaring any use of any of those words or letters in any order to be a transformation of that work without any other context or evidence of transformation is silly. We could rephrase the OP's post as:

>I found [logs of users from Paramount's writers offices reading] my blog recently. Assuming they used that data, when will they attribute me?

To see that the idea on the face is silly. OP has no evidence that any of their work was used at all, or even that what was used could even be covered under the license in the first place.

Re: More writers sue OpenAI for copyright infringement over AI training

#43
post #26

Earlier quoted context omitted.

Like what? The only things I can think of relate to quality/safety (eg. drivers or lawyers).

Participate politically, seek employment, and of course the examples you gave. When pubic safety and goodwill comes in to focus, that's where the role of automation is scrutinized and minimized more heavily. Copyright itself is an invention and area of balancing individual rights and greater public good. Machines are not human and they are not sentiment and sapient at a level where we can view them differently. Perha…

Human or not doesn't seem relevant in any jurisdiction which emphasizes copyright as a means to the end of "promoting useful art", as the US Constitution does, rather than an end in itself.

The moral calculus probably changes if machines are deemed capable of producing "useful art", as granting artists temporary monopoly ceases to become the only mechanism of spurring that art.

Re: More writers sue OpenAI for copyright infringement over AI training

#44
post #12

Random thought: my blog is licensed under a Creative Commons license [1] that allows you to use and transform my content as long as you give attribution and distribute your contributions under the same terms. I found the OpenAI bot scraping my blog recently. Assuming they used that data, when will they attribute me? [1] https://creativecommons.org/licenses/by-sa/4.0/

Interesting question, continuing on this, since they probably used GPL-3 code with the Affero clause, do they have to open source GPT? (The Affero clause is I believe the more directly applicable license thingy, though CC by-sa should also work.) https://www.gnu.org/licenses/agpl-3.0.en.html

I think all the code license question does not matter much, because the code is data input, not a part of their actual program

Like githubs servers host AGPL code as data, without having to be open-source

The perceived problem there, is if their model generates an exact copy of some AGPL code, and you use it in your project unknowingly, and then you get can sued

Re: More writers sue OpenAI for copyright infringement over AI training

#45

Earlier quoted context omitted.

> Assuming they used that data... That's the key part. You haven't yet proved they have actually used your content for anything (other than, potentially, read the license to decide if they should include or discard from their training set). But in practice we'll never know for sure if they are respecting the terms of licenses until 1) this is tested in court, or 2) there's some internal leak that points into either d…

I expect that OpenAI would concede that they used the data in any court case immediately to get that issue off the table, I really don't think they have a strong interest in foot-dragging on this stuff, right? I would think OpenAI wants the thornier legal issues actually settled so that the whole ecosystem can grow within those terms & they can lobby for the legal changes they need/want?

> wants the thornier legal issues actually settled

.. wants the thornier issues to be debated and re-tried ad infinitum, as long as they generate cash flow and build their moat(s).. more likely

Re: More writers sue OpenAI for copyright infringement over AI training

#46
post #9

Earlier quoted context omitted.

If you can ask ChatGPT about any book contents, you don't need to get the book, and if you don't need to get the book then author got robbed, ClosedAI/MS profited.

This reminds me of the tenuous RIAA claim that every pirated piece of media represented a lost sale back when they were suing their customers in the 2000s.

It's been something of a wild ride for me having lived through the "Information wants to be free" era to now live in the new "Reading my publicly published writings and deriving new things from that is theft" era. The next few years of court battles around this are going to be interesting, and I'm not too hopeful on the odds that the "little guy" wins in the end. Seemingly "little guy" affirming results might just turn around and further entrench large players instead.

Re: More writers sue OpenAI for copyright infringement over AI training

#47

Earlier quoted context omitted.

If you can ask ChatGPT about any book contents, you don't need to get the book, and if you don't need to get the book then author got robbed, ClosedAI/MS profited.

I think it's important to distinguish between content and presentation. Most books don't offer entirely new content, but (at best) give some novel way of presenting old content. Consider a modern retelling of Greek Mythology. The stories weren't the original contribution by the author (but by the Ancient Greeks), but the particular way they tell it may be. So ChatGPT telling people about its "content" is unproblemati…

it appears from your emphasis that you are arguing generally that "originality" and personal authorship are rare in practice, and therefore imply that mixing in training is "mostly not infringement"

I disagree with this emphasis, given that rote, repetitive or technical material that is not original authorship is not in peril. Human authors who wrote original creative content, or wrote in a style that is personal and widely recognized, their rights to trade and commerce are in peril. That is much more important over the long term, and is not worth losing for convenient information mixers.

Re: More writers sue OpenAI for copyright infringement over AI training

#48

this is a great lawsuit! if you read the complaint, they catch OpenAI dead-to-rights .. asking about plot details with names from the books, asking to write a paragraph in the style of the author in that book, and a diversity of authors that shows social awareness.. great support for this from California

If you go to wikipedia and look up a book, you'll likely find plenty of plot details, including character names. Is this also infringing?

As far as style goes, copyright doesn't protect that. Trademark MIGHT if your style is distinctive enough to be a trademark (and is used as such), but the "style" of a writer is largely about tempo and word choices, none of which are subject to copyright protections.

Re: More writers sue OpenAI for copyright infringement over AI training

#49

Earlier quoted context omitted.

If you can ask ChatGPT about any book contents, you don't need to get the book, and if you don't need to get the book then author got robbed, ClosedAI/MS profited.

If you ask a human questions about a book, thereby avoiding having to buy the book for class, did you rob the author?

is a sale forced or coerced, also comes to mind. Tales of college undergrads forced to buy hundreds of dollars worth of books for single semester come to mind...

but let's be direct - are we talking about market share in the millions of views, where pirate copies are also available, or the sale of any books at all compared to a few hundred over a year. Quite the difference on a subsistence level of an individual author, no?

Re: More writers sue OpenAI for copyright infringement over AI training

#50
post #48

this is a great lawsuit! if you read the complaint, they catch OpenAI dead-to-rights .. asking about plot details with names from the books, asking to write a paragraph in the style of the author in that book, and a diversity of authors that shows social awareness.. great support for this from California

If you go to wikipedia and look up a book, you'll likely find plenty of plot details, including character names. Is this also infringing? As far as style goes, copyright doesn't protect that. Trademark MIGHT if your style is distinctive enough to be a trademark (and is used as such), but the "style" of a writer is largely about tempo and word choices, none of which are subject to copyright protections.

I think we are now reproducing multiple generations of debate on this topic, in a few go-rounds.. Let's note that among the four largest economies in the world, they each have different rules for this.
Post reply on HN