Live data from Hacker News

New York Times considers legal action against OpenAI as copyright tensions swirl

npr.org

51–60 of 383 posts

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#53
post #44

Earlier quoted context omitted.

That copy didn’t go into a human’s brain. It went into GPU memory. It’s a copy under the law, no different from copying a Taylor Swift mp3 onto a flash drive. Whether that copy was fair use is the key question.

If I buy a license to listen to a Taylor Swift mp3, I can copy it onto a flash drive (or anywhere I like to use it). And that’s fine until I distribute copies to others.

Yes, and if OpenAI had a license to use the Times’ works as model training material this would be a non-issue, of course.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#54
post #36
post #9

Earlier quoted context omitted.

Paraphrasing is not the issue. The issue is that OpenAI copied the Times ’ creative works into a GPU to train a model. That copy was likely neither licensed nor fair use.

Do you think OpenAI doesn’t subscribe? I too load creative works into all sorts of temporary structures in order to read the paper. I don’t need to license it, I pay for a subscription. Should I pay more if I memorize the paper? Should I pay more if I read it to my sick friend in the hospital? Should I pay more if I save copies to my own hard drive and grep for words in the files? Should robots have a higher subscrip…

Are you using this knowledge to produce a product that materially undercuts NYT revenue?

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#55
post #49

I don't think it's an exaggeration to say that LLMs might lead to the end of the open web, or at least a drastically reduced version of it. So much of these model's utility is in directly competing with the producers of the training data. Content creators and aggregators are seeing more and more reason to restrict and limit access, to avoid having AI companies consume all of their data and then be the ones making mon…

No it's the death of the corporate content hosting web - the open web was never about making money with your blog post/irc chat/usenet group/etc content, at least in my opinion.

Let data be free! If someone wants to use it to make money, well, it's open, just like open source. It's still not okay to take open source work and claim it as your own, which is what copyright should be limited to. Stealing a photo or plagiarizing an essay is intrinsically different than just having a copy read by something, be it human or an mechanical process such as training a LLM.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#56

>If, when someone searches online, they are served a paragraph-long answer from an AI tool that refashions reporting from The Times, the need to visit the publisher's website is greatly diminished, said one person involved in the talks. If, when someone reads a newspaper, they are served a paragraph-long answer from an NYTimes reporter that refashions reporting from local sources, the need to interact with the local…

So? The local source is free to sue the NYT for copyright infringement if they so wish.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#57
post #18

Earlier quoted context omitted.

If I read it and memorize it, my brain has made a copy.

Great, can I offer you as a service to billions of people?

Sorry, but I don't expose my endpoints to just anyone.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#58
post #18
post #9

Earlier quoted context omitted.

Paraphrasing is not the issue. The issue is that OpenAI copied the Times ’ creative works into a GPU to train a model. That copy was likely neither licensed nor fair use.

If I read it and memorize it, my brain has made a copy.

somehow highly doubt "my brain makes copies too so copyright law is invalid" won't clear the legal bar for invalidating their lawsuit.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#59

I think all OpenAI needs to do is scan physical newspapers and OCR them. No ToS to agree to, and no ToS on print editions.

ignorance of copyright law won't save you here. i can't legally torrent a copywrited music file just because "no tos to agree to" when i listen to it.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#60
post #36
post #9

Earlier quoted context omitted.

Paraphrasing is not the issue. The issue is that OpenAI copied the Times ’ creative works into a GPU to train a model. That copy was likely neither licensed nor fair use.

Do you think OpenAI doesn’t subscribe? I too load creative works into all sorts of temporary structures in order to read the paper. I don’t need to license it, I pay for a subscription. Should I pay more if I memorize the paper? Should I pay more if I read it to my sick friend in the hospital? Should I pay more if I save copies to my own hard drive and grep for words in the files? Should robots have a higher subscrip…

a subscription doesn't give you an automatic escape hatch out of copyright law.

here's their ToS, which is pretty clear about what you cannot do: https://help.nytimes.com/hc/en-us/articles/115014893428-Term... (relevant parts below)

Without NYT’s prior written consent, you shall not:

...

(2) use robots, spiders, scripts, service, software or any manual or automatic device, tool, or process designed to data mine or scrape the Content, data or information from the Services, or otherwise use, access, or collect the Content, data or information from the Services using automated means;

(3) use the Content for the development of any software program, including, but not limited to, training a machine learning or artificial intelligence (AI) system.

...

(5) cache or archive the Content (except for a public search engine’s use of spiders for creating search indices);

Post reply on HN