Live data from Hacker News

US gov sides with OpenAI on issue of training LLMs on copyrighted material

techcrunch.com

11–20 of 31 posts

Re: US gov sides with OpenAI on issue of training LLMs on copyrighted material

#11

What's mine is yours, comrade!!!!

The CCP adopted capitalism too. There's no sense in worrying how closely the university aligns real policy with various theoretical economic ideals. If we allow copyright law to directly inhibit technological advancement then we defeat the very purpose of copyright law. That said, it need not inhibit if they just pay for using it. Hopefully they will sue for reasonable compensation, noting that they have such a privi…

Yes, just give me my food ration comrade. It is all I need to keep making art for the glory of The Party and it's Supreme Leader. Everything I do I do for the glory of AI. AI is progress!

Re: US gov sides with OpenAI on issue of training LLMs on copyrighted material

#14
post #9

I think it makes sense to say training on material you legally acquired is fair use. Copyright, quite famously, doesn't protect ideas. Nobody really needs or wants AI models to reproduce verbatim copies of books or images or whatever, and they try not to do this anyway, and just because you could maybe make an image or whatever with a copyrighted character design or something, normal intellectual property law already…

Here's the bottom line for me, take two llms, train one on copyrighted materials, do not train the other at all. now, I'm wondering here, which llm will be more sellable, more able to answer questions etc?

How about those music AIs, why aren't they training on classical music and theory textbooks? Why are they being accused of training on copyrighted music? I remember one chat about music AIs, the proponent/enthusiast mentioned "creating" a song by using "johnny cash" in the prompt to get a song in that distinctive style.

I will be the first to admit how useful these AIs are, while I don't use them to "create music", I sure as heck have used them for chats and to write programs. I'm not one of those naysayers talking about the outputs being crap because despite some errors here and there, I've found great utility from these things and I don't even mess with "frontier models" from openai or anthropic. My gripes are about the costs about what we're doing here.

When chatting about "fair use", we can go with the legal definitions which are clear as mud, but I think a better and more sensible path forward would be to consider all the ramifications of their use. Even the term "fair use" implies results much better than we're actually seeing, the economic forces in play here do not seem fair at all.

Re: US gov sides with OpenAI on issue of training LLMs on copyrighted material

#15
Chinese companies will train their models on all the data they can get their hands on. We can make USA companies act ethically, but then they would become irrelevant. The "safeguards" they were made to implement already makes them less useful than comparable Chinese models. A few more years on that trajectory and we won't have to worry about them anymore

Re: US gov sides with OpenAI on issue of training LLMs on copyrighted material

#16
post #12

So copyright is now effectively dead ?

As long as you transform copyrighted data into something else then it appears so.

Which will be also interesting for prompts used by the users of these AI companies, because even if those prompts contain copyrighted data, they can be used for training because training is transforming these prompts into something else.

Re: US gov sides with OpenAI on issue of training LLMs on copyrighted material

#18
post #9

I think it makes sense to say training on material you legally acquired is fair use. Copyright, quite famously, doesn't protect ideas. Nobody really needs or wants AI models to reproduce verbatim copies of books or images or whatever, and they try not to do this anyway, and just because you could maybe make an image or whatever with a copyrighted character design or something, normal intellectual property law already…

Here's the bottom line for me, take two llms, train one on copyrighted materials, do not train the other at all. now, I'm wondering here, which llm will be more sellable, more able to answer questions etc? How about those music AIs, why aren't they training on classical music and theory textbooks? Why are they being accused of training on copyrighted music? I remember one chat about music AIs, the proponent/enthusias…

I don't understand your point.

Re: US gov sides with OpenAI on issue of training LLMs on copyrighted material

#19
post #18

Earlier quoted context omitted.

Here's the bottom line for me, take two llms, train one on copyrighted materials, do not train the other at all. now, I'm wondering here, which llm will be more sellable, more able to answer questions etc? How about those music AIs, why aren't they training on classical music and theory textbooks? Why are they being accused of training on copyrighted music? I remember one chat about music AIs, the proponent/enthusias…

I don't understand your point.

Its my stated disagreement with the following:

>I think it makes sense to say training on material you legally acquired is fair use.

First an aside, there's no law against accessing copyrighted material. OpenAI is certainly welcome to read the New York Times. I can legally buy dvds but the legality of redistributing rips is only considered should I redistribute them.

One of the considerations of "fair use" might be the possibility of benefits or drawbacks to society of those uses that fall under "fair use" exceptions. I would hold the AI's use to that standard, and thats really when we should take the entirety into consideration. This is a big topic tho. We want laws because we believe they benefit society, when loopholes appear we'd normally like them closed, obviously there are branches of the US government that are simply not "normal" right now so theres that lol

I will also admit that a narrower view of "fair use" is to equate "training" of these AIs with human use of the material, after all someone reading the new york times are certainly not infringing on anybody's copyright, in fact they're likely fufilling the new york times internally held purpose, ie., they do the writing, they hope to be read with whatever profit to them that might bring. Just like an AI reading that page by controlling a web browser right? Well this is a good rabbit hole too and you're welcome to try to present that case also, because I've been of the opinion for the last few years that LLMs are not people, and I can chat about that all day.

Its easy to post that you don't understand, so if you'd like to post that again this is fine with me, I'm often a very misunderstood individual :)

Re: US gov sides with OpenAI on issue of training LLMs on copyrighted material

#20

What's mine is yours, comrade!!!!

The CCP adopted capitalism too. There's no sense in worrying how closely the university aligns real policy with various theoretical economic ideals. If we allow copyright law to directly inhibit technological advancement then we defeat the very purpose of copyright law. That said, it need not inhibit if they just pay for using it. Hopefully they will sue for reasonable compensation, noting that they have such a privi…

Copyright law can be misused and can inhibit some advancement, but the fact that “certain groups” or “certain companies” can just ignore laws with no ramifications is something else. If it’s probably life-saving or a pipeline of a dream (see Elon musk always promising FSD), then they should pay or prove the end result.
Post reply on HN