Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

31–40 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#31

I read a NYT article and publish a summary of facts that I learned: totally legit. Train a model on NYT text that outputs a summary of facts that it learned: OMG literally murder.

That's why there will be a legalization of the fair use. Just let your intellectual to be used for free training material is not sustainable.

Also remember copyright laws was not there in the first place.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#32

Earlier quoted context omitted.

Wow they want to kill it. I wonder if we've just lived through the golden Napster era of LLMs.

Just train on NYT articles no longer in copyright. We may be better for it.

Next thing you know ChatGPT gives you the best way to crank your automobile and take good care of your crinoline.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#33

I've been arguing since ChatGPT came out that LLMs should fall under fair use as a "transformative work". I'm not a lawyer and this is just my non-expert opinion, but it will be interesting to see what the legal system has to say about this.

Suit claims that GPT reproduced passages from NYT almost verbatim.

Precisely.

This tired 'fair use' excuses from AI bros whilst the GPT has reproduced the article text verbatim, word for word and it being monetized without the permission from the copyright holder and source (NYT) is an obvious copyright violation 101. Full stop.

Again, just like Getty v. Stability, this copyright lawsuit will end in a licensing deal. Apple played it smart with their choice with licensing deals to train their GPT [0]. But this time, OpenAI knew they could get a license to train on NYT articles but chose not to.

[0] https://9to5mac.com/2023/12/22/apple-wants-to-train-its-ai-w...

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#34

Companies that have content all see dollar signs. NYT won't mind if you use their content to train LLMs - as long as they get a commission. Reddit will shut down their free API and make you pay to get training content. Discord is going to be selling content for AI training too - if they haven't already done so. Twitter is doing it. They didn't care before because LLMs were just experiments. Now we're talking trillion…

"They" also include the people working there. Why someone work with full time writing articles should give the work for free just let someone to train it and make money out of it as a consequence?

>Why someone work with full time writing articles should give the work for free

They are not giving it out "for free", in fact they're being paid by their employer to write these articles. Moreover, the writers themselves stand noth' to gain from their past writings financially as they don't belong to the ownership structure of the business.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#35

I read a NYT article and publish a summary of facts that I learned: totally legit. Train a model on NYT text that outputs a summary of facts that it learned: OMG literally murder.

Because it's not just summarizing the bare facts. It's a parrot.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#36

Earlier quoted context omitted.

What if they OCR’d the newspapers? No ToS there.

I’m pretty sure there is still a copyright also for the physical newspaper.

For the paper or the author? What exactly was the licensing agreement for Op-Ed authors in 1962?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#37

I've been arguing since ChatGPT came out that LLMs should fall under fair use as a "transformative work". I'm not a lawyer and this is just my non-expert opinion, but it will be interesting to see what the legal system has to say about this.

It's inevitable that this question ends up at the supreme court. And the sooner the better IMO. It's clearly fair use. Generative agents will be seen legally as no different than a human artist leveraging the summation of their influences to create a new work.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#38
post #33

Earlier quoted context omitted.

Suit claims that GPT reproduced passages from NYT almost verbatim.

Precisely. This tired 'fair use' excuses from AI bros whilst the GPT has reproduced the article text verbatim, word for word and it being monetized without the permission from the copyright holder and source (NYT) is an obvious copyright violation 101. Full stop. Again, just like Getty v. Stability, this copyright lawsuit will end in a licensing deal. Apple played it smart with their choice with licensing deals to tr…

The four factors considered in a fair use test:

    the purpose and character of the use
    the nature of the copyrighted work
    the amount and substantiality of the portion taken
    the effect of the use upon the potential market.
Literally every single one of these factors has very complicated precedent and each one is an open question when it comes to AI. Since fair use is a balancing test this could go any way.

Stability took the easy way out because they didn't have billions of dollars to play around with and Microsoft to back them. Let's see what OpenAI does but calling everyone who disagrees with your naive interpretation of fair use "AI bros" is doing everyone a disservice.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#39

I read a NYT article and publish a summary of facts that I learned: totally legit. Train a model on NYT text that outputs a summary of facts that it learned: OMG literally murder.

Fair use is intended for humans, much like copyright in general.

If you can't copyright AI-generated pieces, then why would fair use apply to LLMs?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#40
this was predicted in the very influential epic 2014 video in 02004

https://www.youtube.com/watch?v=eUHBPuHS-7s (the original is flash and has thus been consigned to the memory hole, so we are left with this poor-quality conversion)

36": 'however, the press as you know it has ceased to exist'

40": '20th-century news organizations are an afterthought; a lonely remnant of a not-too-distant past'

2'11": 'also in 2002, google launches google news, a news portal. news organizations cry foul. google news is edited entirely by computers'

5'13": 'the news wars of 2010 are notable for the fact that no actual news organizations take part. googlezon finally checkmates microsoft with a feature the software giant cannot match: using a new algorithm, googlezon's computers construct new stories, dynamically stripping sentences and facts from all content sources, and recombining them. the computer writes a new story for every user'

5'55": 'in 2011 the slumbering fourth estate awakes to make its first and final stand. the new york times company sues googlezon, claiming that the company's fact-stripping robots are a violation of copyright law. the case goes all the way to the supreme court'

they didn't get the details exactly right, but overall the accuracy is astounding

however, that may be a hyperstition artifact in this timeline

https://en.wikipedia.org/wiki/EPIC_2014 (i thought epic 2014 might be the only flash video to hae a wikipedia article about it, but then i looked and found five others)

Post reply on HN