Live data from Hacker News

New York Times considers legal action against OpenAI as copyright tensions swirl

npr.org

341–350 of 383 posts

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#341
post #253

Earlier quoted context omitted.

Short of a police state, how would you enforce this? This has napster -> subscription spotify energy. But the only people happy about that are Spotify and people who found it distasteful to download music illegally. There just wasn’t a consumer-friendly option for a while, so the black market was the only market. So. The enforcement mechanism is what… a scary DMCA letter? (There will definitely be a stupid DCAIA in t…

The copyright holder gets a share of ownership in any AI model derived from its work, and thus a share of any resulting revenue.

All 10 million of them?

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#342

Honestly, I think generative AI losing a massive copyright showdown is inevitable at this stage. It's extremely easy to get the latest generation of AIs to produce outputs that in many fields sans-AI would be trivially considered as IP infringement. While there are many interesting reasonable legal & technical arguments that it's not, the result completely undermines copyright protections regardless. If that's accept…

I will be interested to see how the copyright suit plays out with Taylor Swift vs. Guy who asked an AI to make a new song that sounds like a generic Taylor Swift song.

[deleted]

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#343

Earlier quoted context omitted.

Ruling in favor of copyright will call into question search engines and the like as well. Do you think Bing or Google are going to negotiate copying rights with the world's websites? LLMs are proving that intellectual property has a bunch of holes in it. It's been unstable ground to defend since day one. Upon what principle should we believe that one can own an idea and all performances or derivatives of it? Patents…

I think there is a major qualitative difference between generative AI and search engines. Search engines index the web and point you at other people's work, along the way showing perhaps too much of that content (thus "stealing" users from the target webpage). But they don't reshuffle existing content into something apparently new and original. The "malicious" case for generative AI is that it sucks in copyrighted wo…

Indexing the copyrighted works is rehashing it. You're rehashing it into a different format that is more easily searchable by a computer. But that work is still based on other people's copyrighted works.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#344
post #341

Earlier quoted context omitted.

The copyright holder gets a share of ownership in any AI model derived from its work, and thus a share of any resulting revenue.

All 10 million of them?

Any large corporation likely has more individual shareholders than that (particularly when you include indirect ownership via mutual funds).

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#345

Earlier quoted context omitted.

Ruling in favor of copyright will call into question search engines and the like as well. Do you think Bing or Google are going to negotiate copying rights with the world's websites? LLMs are proving that intellectual property has a bunch of holes in it. It's been unstable ground to defend since day one. Upon what principle should we believe that one can own an idea and all performances or derivatives of it? Patents…

I think there is a major qualitative difference between generative AI and search engines. Search engines index the web and point you at other people's work, along the way showing perhaps too much of that content (thus "stealing" users from the target webpage). But they don't reshuffle existing content into something apparently new and original. The "malicious" case for generative AI is that it sucks in copyrighted wo…

I'm not saying search engines will be considered the same as LLMs, but given that LLMs are pushing the tolerances of Fair Use, should that get limited through litigation or legislation, the things search engines get away with like caching or summarizing, AMP pages, etc may cease to be legal and search engines may have to adopt less rich means of communicating relevancy, perhaps by showing the keywords that match or something else indirect but still true about a source, without pulling content straight from it.

As it stands right now, yes, Fair Use. But where do we draw the line? That line's been blurry for a while. Mostly limited to no more than 30 seconds of a performance, and no more than what's needed to quote literary works. Not sure about lyrics.

The issue is on some level, most things we create are derivative. Someone had to have the idea first, but once an idea is unleashed upon the world, it seems very difficult and unwieldy to put the genie back in the bottle.

I'm not really pro-LLM since it enables business to leech off of FOSS even more efficiently. The disruption of LLMs seems to be accelerating our philosophical re-examining of copyright and licensing terms in FOSS. We will inevitably need a GPL in the future that disallows remixing via LLM or other generative text algos, due mostly because ensuring all of a result is freely usable is not easy. Limitations could be built in to only 'fetch' code licensed under permissive terms, like MIT or BSD, but the tendency of these models to 'hallucinate' means you really cannot know the legal standing of LLM-generated code, at present.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#346
post #307

Earlier quoted context omitted.

Any form of lossy compression is an irreversible transformation. We do it all the time for video, audio and images (you can't recover the original data) and they are still copyrighted

when you compress a video, it doesn't recreate a new movie with a different story, different lines of text, different scenes and a different compositions for scenes that are similar to the "orginial".

But what is being compressed is the entire corpus of text. It's compressed into model weights. It's the weights that might be under copyright of the authors of the texts that trained it.

The weights are also executable code (in some sense). When you query an LLM you're running this program with a given input. Yeah when it runs it tells a whole lot of things (sometimes novel combinations, sometimes verbatim repetition of trained data) but the point here isn't whether the output of the LLM is copyrighted; it's the weights.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#348

Honestly, I think generative AI losing a massive copyright showdown is inevitable at this stage. It's extremely easy to get the latest generation of AIs to produce outputs that in many fields sans-AI would be trivially considered as IP infringement. While there are many interesting reasonable legal & technical arguments that it's not, the result completely undermines copyright protections regardless. If that's accept…

Also even if it’s ruled to be infringing, these current models aren’t going to go away. And given the additional high-quality training material this allows, it’s fairly likely these models will have an ongoing advantage in quality of output.

So now you’ve divided the world into those who use the best tech, and those who are not allowed. And that is what openAI wants, they’re betting courts are going to rule that however tainted the source, that we can’t put lightning back into the bottle.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#349

Earlier quoted context omitted.

I don't think that no one would create new software/music/books/movies/art/etc. without copyright. Humans have done so for millenia before, they still did so in absence of copyright protections. I don't see how this is not a serious alternative - the only major losers would be the middlemen, not the artists themselves.

Eh, I could definitely see the artists losing. The most obvious scenario that comes to mind for me is, imagine an independent artists launching their (book/film/album/etc) and the same day someone with more resources and experience takes the work and markets it better than the OG author ever could on their own.

In a copyright-less system you’d monetize the creation (eg via patronage), so someone else distributing the work is fine, even helpful.

It makes no sense to add artificial scarcity to ideas, the cost of replication is inherently zero and all kinds of noxious consequences on society (like derivative works being impossible to create) shake out as a result. If the goal is encouraging creation why are you punishing and penalizing the creation of derivative works?

Letting megacorps monopolize popular culture for 100+ years after its inception is a relatively newfangled idea and it’s baffling that so many people just blindly accept that this is the way it has to be. Copyright works against individuals in almost all cases, we benefit much more from free interchange of ideas. Social diffusion and remixing is a fundamental human force and this AI stuff forced the issue by doing the exact same things with impossible precision and scale, such that the absurdity of the system for humans is revealed as well.

The current copyright regime is the social equivalent of “we could have a cop taking down license plates if we wanted”. And AI does the “so it’s therefore legal to automatically and instantly record everyone’s license plates at every intersection in the country 24/7”. The principle is the same but the ease of use reveals the absurdity of the principle.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#350
post #320

Earlier quoted context omitted.

Does that invalidate the analogy? The point I'm trying to make is if you rule these behaviors illegal, then you're necessarily making intelligent AI illegal, because humans are capable of the same behaviors.

Plagiarism is already a copyright violation so I’m not sure what your point is, and what you’re alluding to is literally what the post is about…

Plagiarism is not a copyright violation, it’s a violation of the social rules of private (academic) institutions and imposed by those institutions upon their members.

If plagiarism was a copyright violation then citing the source wouldn’t make the copyright violation go away. Putting the artists name in the title of the YouTube doesn’t make it not a copyright violation.

Post reply on HN