Live data from Hacker News

More writers sue OpenAI for copyright infringement over AI training

reuters.com

51–60 of 68 posts

Re: More writers sue OpenAI for copyright infringement over AI training

#51

Random thought: my blog is licensed under a Creative Commons license [1] that allows you to use and transform my content as long as you give attribution and distribute your contributions under the same terms. I found the OpenAI bot scraping my blog recently. Assuming they used that data, when will they attribute me? [1] https://creativecommons.org/licenses/by-sa/4.0/

Your license doesn't override copyright law.

Given that Google successfully used a fair use defense in Authors Guild, Inc. v. Google, Inc., I think it's likely OpenAI and the others will also win in court.

I do think it's possible for specific uses of the output of LLMs to be copyright infringement. That's why it's interesting to see Microsoft to indemnify customers of their commercial products in the event a case is brought against the customer. This is smart on Microsoft's part; the risk probably isn't very high and by making it a non-issue for their customers, many more will feel comfortable using their LLM-based features and services.

Re: More writers sue OpenAI for copyright infringement over AI training

#52

Earlier quoted context omitted.

I think it's important to distinguish between content and presentation. Most books don't offer entirely new content, but (at best) give some novel way of presenting old content. Consider a modern retelling of Greek Mythology. The stories weren't the original contribution by the author (but by the Ancient Greeks), but the particular way they tell it may be. So ChatGPT telling people about its "content" is unproblemati…

it appears from your emphasis that you are arguing generally that "originality" and personal authorship are rare in practice, and therefore imply that mixing in training is "mostly not infringement" I disagree with this emphasis, given that rote, repetitive or technical material that is not original authorship is not in peril. Human authors who wrote original creative content, or wrote in a style that is personal and…

You might find this documentary interesting: "Everything Is A Remix" [1]

[1] https://youtu.be/nJPERZDfyWc?si=IooGFXhb5gbYNWyS

Re: More writers sue OpenAI for copyright infringement over AI training

#53
post #5

Earlier quoted context omitted.

Not at all like piracy. When someone pirates a book, they're replacing the original without consent or remuneration to the copyright holders. When you train an AI on the contents of a book, you're not replacing it. If someone is interested in the content, they still need to buy it. Using ChatGPT is not a substitute. If it is, they're gonna have to prove it in court, but I doubt they'll be able to.

If you can ask ChatGPT about any book contents, you don't need to get the book, and if you don't need to get the book then author got robbed, ClosedAI/MS profited.

Are copyright holders being robbed by professors that answer their students' questions?

Don't teachers do the same?

- Trained their minds on existing books

- Tutor the next generation of students

- Give classes on book contents

- Answer questions about those books

The book publishing industry didn't go out of business because there are teachers answering questions. To the contrary, it benefited book sales, because most people aren't good self-learners.

What's wrong with having a machine do the same?

Re: More writers sue OpenAI for copyright infringement over AI training

#54

Earlier quoted context omitted.

I think it's important to distinguish between content and presentation. Most books don't offer entirely new content, but (at best) give some novel way of presenting old content. Consider a modern retelling of Greek Mythology. The stories weren't the original contribution by the author (but by the Ancient Greeks), but the particular way they tell it may be. So ChatGPT telling people about its "content" is unproblemati…

it appears from your emphasis that you are arguing generally that "originality" and personal authorship are rare in practice, and therefore imply that mixing in training is "mostly not infringement" I disagree with this emphasis, given that rote, repetitive or technical material that is not original authorship is not in peril. Human authors who wrote original creative content, or wrote in a style that is personal and…

> Human authors who wrote original creative content, or wrote in a style that is personal and widely recognized, their rights to trade and commerce are in peril

I see what you're saying, but I fail to see how ChatGPT merely copying their style (not: content) might impact "their rights to trade and commerce". Suppose I ask ChatGPT to "tell me some jokes in the style of Louis CK". Would that make me less likely to stream a Louis CK comedy special?

(By contrast, if I ask ChatGPT to summarize the key revelations from a book like Fire and Fury, that probably would make me less likely to buy the book, because if I buy the book it'd be for the novel information contained in it, but ChatGPT already divulged it to me.)

Re: More writers sue OpenAI for copyright infringement over AI training

#55
post #48

Earlier quoted context omitted.

If you go to wikipedia and look up a book, you'll likely find plenty of plot details, including character names. Is this also infringing? As far as style goes, copyright doesn't protect that. Trademark MIGHT if your style is distinctive enough to be a trademark (and is used as such), but the "style" of a writer is largely about tempo and word choices, none of which are subject to copyright protections.

I think we are now reproducing multiple generations of debate on this topic, in a few go-rounds.. Let's note that among the four largest economies in the world, they each have different rules for this.

Do any of those economies really have laws protecting an author's "style"? Because I'd really like to see a legal definition of an author's style, and a case that found someone guilty of infringing on that style (separate from trademark and copyright of course)

Re: More writers sue OpenAI for copyright infringement over AI training

#56

Earlier quoted context omitted.

I expect that OpenAI would concede that they used the data in any court case immediately to get that issue off the table, I really don't think they have a strong interest in foot-dragging on this stuff, right? I would think OpenAI wants the thornier legal issues actually settled so that the whole ecosystem can grow within those terms & they can lobby for the legal changes they need/want?

> wants the thornier legal issues actually settled .. wants the thornier issues to be debated and re-tried ad infinitum, as long as they generate cash flow and build their moat(s).. more likely

https://the-decoder.com/openai-apparently-going-all-in-on-ch...

This behaviour seems more consistent with wanting is sorted out than stalling for time.

Re: More writers sue OpenAI for copyright infringement over AI training

#58

Random thought: my blog is licensed under a Creative Commons license [1] that allows you to use and transform my content as long as you give attribution and distribute your contributions under the same terms. I found the OpenAI bot scraping my blog recently. Assuming they used that data, when will they attribute me? [1] https://creativecommons.org/licenses/by-sa/4.0/

How recent? Because ChatGPT is always on the same mantra, of its training being from back September 2021 with no updates...Even for ChatGPT-4

Re: More writers sue OpenAI for copyright infringement over AI training

#59

Earlier quoted context omitted.

> wants the thornier legal issues actually settled .. wants the thornier issues to be debated and re-tried ad infinitum, as long as they generate cash flow and build their moat(s).. more likely

https://the-decoder.com/openai-apparently-going-all-in-on-ch... This behaviour seems more consistent with wanting is sorted out than stalling for time.

a motion to dismiss is "going all in" ?
Post reply on HN