Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

41–50 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#41
The suit demonstrates instances where ChatGTP / Bing Copilot copy from the NYT verbatim. I think it is hard to argue that such copying constitutes "fair use". However, OAI/MS should be able to fix this within the current paradigm: Just learn to recognize and punish plagiarism via RLHF.

However, the suit goes far beyond claiming that such copying violates their copyright: "Unauthorized copying of Times Works without payment to train LLMs is a substitutive use that is not justified by any transformative purpose."

This is a strong claim that just downloading articles into training data is what violates the copyright. That GTP outputs verbatim copies is a red herring. Hopefully the judge(s) will notice and direct focus on the interesting, high-stakes, and murky legal issues raised when we ask: What about a model can (or can't) be "transformative"?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#42
post #33

Earlier quoted context omitted.

Suit claims that GPT reproduced passages from NYT almost verbatim.

Precisely. This tired 'fair use' excuses from AI bros whilst the GPT has reproduced the article text verbatim, word for word and it being monetized without the permission from the copyright holder and source (NYT) is an obvious copyright violation 101. Full stop. Again, just like Getty v. Stability, this copyright lawsuit will end in a licensing deal. Apple played it smart with their choice with licensing deals to tr…

> AI bros

What (or whom) do you consider to be an "AI bro?"

This sort of ad hominem generalization usually accompanies a weak argument.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#43
post #33

Earlier quoted context omitted.

Precisely. This tired 'fair use' excuses from AI bros whilst the GPT has reproduced the article text verbatim, word for word and it being monetized without the permission from the copyright holder and source (NYT) is an obvious copyright violation 101. Full stop. Again, just like Getty v. Stability, this copyright lawsuit will end in a licensing deal. Apple played it smart with their choice with licensing deals to tr…

> AI bros What (or whom) do you consider to be an "AI bro?" This sort of ad hominem generalization usually accompanies a weak argument.

It seems to be used by people who've previously used the term "tech bro."

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#44

I've been arguing since ChatGPT came out that LLMs should fall under fair use as a "transformative work". I'm not a lawyer and this is just my non-expert opinion, but it will be interesting to see what the legal system has to say about this.

Suit claims that GPT reproduced passages from NYT almost verbatim.

I don’t doubt it does. It’s easy to get it to spit out long answers from Stack Overflow verbatim, I’ve done it. Maybe some of the “transformative” nature of the LLM output is the removal of any authorship, copyright, license, and edit history information. ;) The point here is to supplant Google as the portal of information, right? It doesn’t have new information, but it’s pretty good at remixing the words from multiple sources, when it has multiple sources. One possible reason for their legal woes wrt copyright is that it’s also great at memorizing things that only have one source. My college Markov-chain text predictor would do the same thing and easily get stuck in local regions if it couldn’t match something else.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#45

Companies that have content all see dollar signs. NYT won't mind if you use their content to train LLMs - as long as they get a commission. Reddit will shut down their free API and make you pay to get training content. Discord is going to be selling content for AI training too - if they haven't already done so. Twitter is doing it. They didn't care before because LLMs were just experiments. Now we're talking trillion…

> They didn't care before because LLMs were just experiments. Now we're talking trillions of dollars of value. Can you make the argument this was their fault for not having forward vision/being asleep at the wheel and "accidentally, in hindsight" letting OpenAI/others have free, open, unlimited access to their content?

Basically none of the training material for GPT was used under an "unlimited" license. There are very important legal limitations. GPT just doesn't care much about them.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#46
post #33

Earlier quoted context omitted.

Precisely. This tired 'fair use' excuses from AI bros whilst the GPT has reproduced the article text verbatim, word for word and it being monetized without the permission from the copyright holder and source (NYT) is an obvious copyright violation 101. Full stop. Again, just like Getty v. Stability, this copyright lawsuit will end in a licensing deal. Apple played it smart with their choice with licensing deals to tr…

> AI bros What (or whom) do you consider to be an "AI bro?" This sort of ad hominem generalization usually accompanies a weak argument.

Not saying I agree with this labeling, but it means approximately the same thing as “crypto bro”, but for AI

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#47

Seems reasonable - they probably broke the TOS of the site

Did OpenAI agree to those ToS? If not, I think (IANAL) LinkedIn was kind enough to give precedent that it's irrelevant.

( https://en.wikipedia.org/wiki/HiQ_Labs_v._LinkedIn )

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#48
post #34

Earlier quoted context omitted.

"They" also include the people working there. Why someone work with full time writing articles should give the work for free just let someone to train it and make money out of it as a consequence?

>Why someone work with full time writing articles should give the work for free They are not giving it out "for free", in fact they're being paid by their employer to write these articles. Moreover, the writers themselves stand noth' to gain from their past writings financially as they don't belong to the ownership structure of the business.

> the writers themselves stand noth' to gain from their past writings financially as they don't belong to the ownership structure of the business.

This is a dumb argument. We're not just talking about ancient articles. We're talking about new content, including content that is yet to be written.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#49
Won't hold in court. GPT is a platform mainly providing answer to private individuals asking. Is like you ask a professor a question and he answered verbatim what copyrighted materials available (due to photographic memory) word for word back to you. Now if you take this answer and write a book or publish enmass on blogs for example, then you are the one should be sued by NYT. If GPT use the exact same wordings and publish it out to evetyone visiting their page, then that is on OpenAI.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#50

Earlier quoted context omitted.

I’m pretty sure there is still a copyright also for the physical newspaper.

For the paper or the author? What exactly was the licensing agreement for Op-Ed authors in 1962?

Read the article. It's not difficult to get ChatGPT to regurgitate recent, obviously copyrighted articles, verbatim.
Post reply on HN