I think all OpenAI needs to do is scan physical newspapers and OCR them. No ToS to agree to, and no ToS on print editions.
ignorance of copyright law won't save you here. i can't legally torrent a copywrited music file just because "no tos to agree to" when i listen to it.
New York Times considers legal action against OpenAI as copyright tensions swirl
61–70 of 383 posts
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#62Earlier quoted context omitted.
If I read it and memorize it, my brain has made a copy.
somehow highly doubt "my brain makes copies too so copyright law is invalid" won't clear the legal bar for invalidating their lawsuit.
Anyway, I've always been a bit prickly about IP stuff ever since a troll lawyer threatened me over a TI-BASIC game when I was like 14. I'm also sure I'm completely wrong-headed about this whole thing and overly anthropomorphizing the LLM.
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#63Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#64Earlier quoted context omitted.
Do you think OpenAI doesn’t subscribe? I too load creative works into all sorts of temporary structures in order to read the paper. I don’t need to license it, I pay for a subscription. Should I pay more if I memorize the paper? Should I pay more if I read it to my sick friend in the hospital? Should I pay more if I save copies to my own hard drive and grep for words in the files? Should robots have a higher subscrip…
a subscription doesn't give you an automatic escape hatch out of copyright law. here's their ToS, which is pretty clear about what you cannot do: https://help.nytimes.com/hc/en-us/articles/115014893428-Term... (relevant parts below) Without NYT’s prior written consent, you shall not: ... (2) use robots, spiders, scripts, service, software or any manual or automatic device, tool, or process designed to data mine or sc…
As an example, my local library offers full-text NYT articles through both nytimes.com and ProQuest.
Notably, ProQuest's terms only explicitly ban scraping metadata and developing software or services that "compete or interfere" with ProQuest products:
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#65Earlier quoted context omitted.
Did the Times grant a license to every router on the internet to transmit its intellectual property to other routers? If not, the judge should grant an injunction contingent on requiring the Times to verify that every person who accesses their content is doing so only over routers and other devices with express written authorization, for every step in the process. Maybe even extend it to browsers and client libraries…
OpenAI is pretty clearly using their work to make derivative content that in certain cases (CNET) is a direct competitor. Honestly, this seems open and shut
If I read five calculus textbooks and write a new one, I don't think that's derivative content (or maybe it is?) Seems like that's what an LLM does - read many works, write a new work.
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#66Earlier quoted context omitted.
Right. But the comment above said the issue is with the copy made for training. The "read it and memorized it" copy, not the "write it back down" copy.
That copy didn’t go into a human’s brain. It went into GPU memory. It’s a copy under the law, no different from copying a Taylor Swift mp3 onto a flash drive. Whether that copy was fair use is the key question.
But as you point out in the router case above, it's transient.
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#67Earlier quoted context omitted.
a subscription doesn't give you an automatic escape hatch out of copyright law. here's their ToS, which is pretty clear about what you cannot do: https://help.nytimes.com/hc/en-us/articles/115014893428-Term... (relevant parts below) Without NYT’s prior written consent, you shall not: ... (2) use robots, spiders, scripts, service, software or any manual or automatic device, tool, or process designed to data mine or sc…
Full text New York Times articles are available through subscription services other than nytimes.com. As an example, my local library offers full-text NYT articles through both nytimes.com and ProQuest. Notably, ProQuest's terms only explicitly ban scraping metadata and developing software or services that "compete or interfere" with ProQuest products: https://about.proquest.com/en/about/terms-and-conditions
> Restrictions. Except as expressly permitted above, Customer and its Authorized Users shall not:
> Remove any copyright and other proprietary notices placed upon the Service or any materials retrieved from the Service by ProQuest or its licensors;
> Perform automated searches against ProQuest’s systems (except for non-burdensome federated search services), including automated “bots,” link checkers or other scripts;
> Provide access to or use of the Services by or for the benefit of any unauthorized school, library, organization, or user;
> Publish, broadcast, sell, use or provide access to the Service or any materials retrieved from the Service in any manner that will infringe the copyright or other proprietary rights of ProQuest or its licensors;
> Download all or parts of the Service in a systematic or regular manner or so as to create a collection of materials comprising all or a material subset of the Service, in any form.
> Store any information on the Service that violates applicable law or the rights of any third party.
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#68Earlier quoted context omitted.
ignorance of copyright law won't save you here. i can't legally torrent a copywrited music file just because "no tos to agree to" when i listen to it.
There is no copyright involved in training a model. No precedent at least. Only online ToS/API restrictions exist for scraping content. The DMCA issues exist entirely on the (re)distribution side of coyright material. So seeding in a torrent swarm = redistribution. If you download some copyright material somehow and don't share it with anyone, there is no caselaw that says anything about it.
> If you download some copyright material somehow and don't share it with anyone, there is no caselaw that says anything about it.
this is still a violation of the law. you cannot download copyrighted material against the terms of the copyright holder (like downloading a movie or album).
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#69I don't think it's an exaggeration to say that LLMs might lead to the end of the open web, or at least a drastically reduced version of it. So much of these model's utility is in directly competing with the producers of the training data. Content creators and aggregators are seeing more and more reason to restrict and limit access, to avoid having AI companies consume all of their data and then be the ones making mon…
No it's the death of the corporate content hosting web - the open web was never about making money with your blog post/irc chat/usenet group/etc content, at least in my opinion. Let data be free! If someone wants to use it to make money, well, it's open, just like open source. It's still not okay to take open source work and claim it as your own, which is what copyright should be limited to. Stealing a photo or plagi…
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#70Moreover, AI would seem to be even more susceptible to capture and manipulation than conventional media.
When it's a question of guiding thought I prefer the humanities to tech. (Same with art.)