Live data from Hacker News

More writers sue OpenAI for copyright infringement over AI training

reuters.com

31–40 of 68 posts

Re: More writers sue OpenAI for copyright infringement over AI training

#31

Random thought: my blog is licensed under a Creative Commons license [1] that allows you to use and transform my content as long as you give attribution and distribute your contributions under the same terms. I found the OpenAI bot scraping my blog recently. Assuming they used that data, when will they attribute me? [1] https://creativecommons.org/licenses/by-sa/4.0/

These AI companies not complying with the licenses on code has meant since Microsoft released their code generator I haven't contributed a single line of open source software nor released any of my projects that way. I removed a bunch a while ago and I will likely remove all of them when I get around to it. I have been fixing bugs and releasing open source projects for decades and I just stopped the moment they did that. Open source is dead to me if the licenses can't be enforced.

Re: More writers sue OpenAI for copyright infringement over AI training

#32

Random thought: my blog is licensed under a Creative Commons license [1] that allows you to use and transform my content as long as you give attribution and distribute your contributions under the same terms. I found the OpenAI bot scraping my blog recently. Assuming they used that data, when will they attribute me? [1] https://creativecommons.org/licenses/by-sa/4.0/

Where's your attributions for all the words written in your comment ;P you remixed the words and grammar patterns from other people's creative common's licenses of other people writings! Note: i'm declaring my comment license as https://creativecommons.org/licenses/by-sa/4.0/ So if you remix or transform my comment by responding it, please attribute to me your response.

Humans are not LLMs trained and operated by a company for profit. Your argument is that LLMs hold all the same basic rights as humans but they hold (and should hold) exactly none.

Re: More writers sue OpenAI for copyright infringement over AI training

#33
post #22

All these lawsuits will die. Why? Because people train on corpuses of data all the time, without a license or any attribution. Every piece of text a writer reads is training that writer. Every image an artist sees helps to train that artist. Every sound a musician hears is training that musician. That doesn't mean they can't exclude their works from training via a license going foreward. But that becomes an enforceme…

LLMs are not humans and your anthropomorphizing argument is idiotic

Re: More writers sue OpenAI for copyright infringement over AI training

#34
post #22

All these lawsuits will die. Why? Because people train on corpuses of data all the time, without a license or any attribution. Every piece of text a writer reads is training that writer. Every image an artist sees helps to train that artist. Every sound a musician hears is training that musician. That doesn't mean they can't exclude their works from training via a license going foreward. But that becomes an enforceme…

People can do a lot of things that we don't legally allow machines or automation to do.

True of some things, but not of Fair Use. Automatically generating thumbnails is generally Fair Use, for example.

Re: More writers sue OpenAI for copyright infringement over AI training

#35
post #5

Earlier quoted context omitted.

Not at all like piracy. When someone pirates a book, they're replacing the original without consent or remuneration to the copyright holders. When you train an AI on the contents of a book, you're not replacing it. If someone is interested in the content, they still need to buy it. Using ChatGPT is not a substitute. If it is, they're gonna have to prove it in court, but I doubt they'll be able to.

If you can ask ChatGPT about any book contents, you don't need to get the book, and if you don't need to get the book then author got robbed, ClosedAI/MS profited.

I think it's important to distinguish between content and presentation. Most books don't offer entirely new content, but (at best) give some novel way of presenting old content. Consider a modern retelling of Greek Mythology. The stories weren't the original contribution by the author (but by the Ancient Greeks), but the particular way they tell it may be. So ChatGPT telling people about its "content" is unproblematic if it's just telling people how the story goes, and only potentially problematic if it's effectively quoting from the book or mimicking its presentation. (And we all know that if ChatGPT is good at one thing, it's paraphrase or re-expressing the same ideas in substantially different ways, so even if ChatGPT literally copies a book's presentation/wording, that would probably have happened by accident rather than necessity)

The vast majority of publications (especially those of a explanatory nature) do not contribute original content/information. The exceptions are things like research articles/monographs, historical records, government reports. But copyright infringement doesn't apply here because these things weren't published with a profit motive but precisely to publicize the information as widely as possible. The only problem area I can think of involves books published by commercial publishers which promise 'exclusive peek' into the life of some famous person (think biographies of celebrities or books like Fire and Fury). In that kind of case there is indeed original content, and revealing it in detail will arguably mean less sales for the authors/publishers.

Re: More writers sue OpenAI for copyright infringement over AI training

#36
post #26

Earlier quoted context omitted.

People can do a lot of things that we don't legally allow machines or automation to do.

Like what? The only things I can think of relate to quality/safety (eg. drivers or lawyers).

Participate politically, seek employment, and of course the examples you gave.

When pubic safety and goodwill comes in to focus, that's where the role of automation is scrutinized and minimized more heavily. Copyright itself is an invention and area of balancing individual rights and greater public good.

Machines are not human and they are not sentiment and sapient at a level where we can view them differently. Perhaps they will change one day, but as it is today these systems are not entitled to do the same things humans get to do. They are tools performing a task, so the laws apply to them as they apply to, well, machines; copying and reproducing whole code blocks or novel chapters without attribution or licenses is something we allow a human to do in their head and not what we allow a machine to do in a prompt, regardless of the non-human mechanisms in between.

Re: More writers sue OpenAI for copyright infringement over AI training

#37
The fair use argument is quite strong

If you dissect the plaintiffs claim they are arbitrarily conflating training and regurgitating

Training is using for criticism and comparison purposes, hence fair use

And there is no lawsuit against what it regurgitates and the purpose of its output, whether someone asks it to give a list for comparison purposes, or specifically asks it for a story that has a plagiarized result

Re: More writers sue OpenAI for copyright infringement over AI training

#39

Random thought: my blog is licensed under a Creative Commons license [1] that allows you to use and transform my content as long as you give attribution and distribute your contributions under the same terms. I found the OpenAI bot scraping my blog recently. Assuming they used that data, when will they attribute me? [1] https://creativecommons.org/licenses/by-sa/4.0/

training will be ruled fair use which doesn't require any license, while there is no lawsuit on the output

Re: More writers sue OpenAI for copyright infringement over AI training

#40

I'm miffed. I tried a couple characters from my books, and zilch: ===== who is dan markunas ChatGPT I'm sorry, but I don't have any information on a person named Dan Markunas in my database .... who is janet saunders ChatGPT I'm sorry, but I don't have any specific information about a person named Janet Saunders in my database, ===========

Your book was published somewhere mid-2021, right?

Two books. The second was after their cutoff date.
Post reply on HN