Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

171–180 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#171
post #164

Earlier quoted context omitted.

Machine learning models have been trained with copyrighted data for a long time. Imagenet is full of copyrighted images, clearview literally just scanned the internet for faces, and I am sure there are other, older examples. I am unsure if this has been tested as fair use by a US court, but I am guessing it will be considered to be so if it is not already.

Excellent, so you're saying I'll be able to download any copyrighted work from any pirate site and be free of all consequence if I just claim that I'm training an AI?

That's not was he's saying at all. He's saying you can train an AI on copyrighted material just like people can learn from copyrighted material.

If you acquire the material illegally that a separate issue that training AI doesn't give you any protection against.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#172
post #170
post #90

Earlier quoted context omitted.

> you must have legally purchased a copy, or been given it by someone who did so (for example, if you received it as a gift). I am also NAL, but can I imagine it goes further than that. Just purchasing a copy doesn't let you create and sell (directly as content or indirectly via a service like a chatbot) derivative works that are substantially similar in style and voice to the original work. For example, an LLM 's re…

> IMO doesn't constitute fair use Yes, the question of whether the way LLMs use the content they use qualifies as fair use is a separate question. My point was simply that that question can't even be reached if the maker of the LLMs doesn't have a legal right to fair use in the first place (because they don't legally own their copy).

> My point was simply that that question can't even be reached if the maker of the LLMs doesn't have a legal right to fair use in the first place (because they don't legally own their copy).

I agree, and I expect that eventually we will start seeing injunctions against creators requiring them to remove content that they don't have legal access to from their training data sets.

And this will probably end up at the Supreme Court.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#173
post #108

Earlier quoted context omitted.

The lawsuit doesn't even mention Google.

No, I did. What's your point?

The point is that GP has no reason to believe Google, like Meta, also used copyrighted materials for training its AI. Why did Sarah Silverman sue OpenAI and Meta but not Google?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#174
If they ripped all of Bibliotik, the more interesting story to me is how they were able to get it all without hitting ratio requirements?

Super fast internet that downloaded all they could before being ratio banned, overwhelmingly fast internet that was hopping on all the popular torrents to slowly build up ratio?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#175

This is actually quite interesting, as it's drawing a distinction between training material that can be accessed by anybody with a web browser (like anybody's blog), vs. training material that was "illegally-acquired... available in bulk via torrent systems." I don't think there's any reason why this would be a relevant legal distinction in terms of distributing an LLM -- blog authors weren't giving consent either. H…

I'm allowed to make private copies of copywritten works. I'm not allowed to redistribute them. To what extent this is redistribution is not clear. Is there much of difference between this model and a machine, like a VCR, that recreates the original work when I press a button?

This is not definitely not redistribution any more than writing a blog post of a book you read is.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#176
post #164

Earlier quoted context omitted.

Excellent, so you're saying I'll be able to download any copyrighted work from any pirate site and be free of all consequence if I just claim that I'm training an AI?

That's not was he's saying at all. He's saying you can train an AI on copyrighted material just like people can learn from copyrighted material. If you acquire the material illegally that a separate issue that training AI doesn't give you any protection against.

But acquiring the material illegally is the thrust of the issue, and the point of the thread here: the notion that large companies can get away with piracy if they just execute it on a large enough scale. Copyright infringment for thee but not for me.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#177

Earlier quoted context omitted.

I can see a good argument in the complaint. The provenance of the training data leads back to it being acquired illegally. Illegally acquired materials were then used in a commercial venture. That the venture was an AI model is perhaps beside the point. You can’t use illegally acquired materials when doing business.

>You can’t use illegally acquired materials when doing business. This vague sentence conjures images of a company building products from stolen parts, but this situation seems different. IANAL, but if I looked at a stolen painting that nobody had ever seen, and sold handwritten descriptions of the painting to whoever wanted to buy one, I'm pretty sure what I've sold is not illegal.

Piracy of content is against the law. All other analogies such as looking at paintings are not at issue here. The content was pirated and there are laws against that, whether we agree with it or not.

So, if the plaintiff can prove the content was pirated, then the use of that content downstream is tainted.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#178
post #50

Earlier quoted context omitted.

Eventually, I imagine a new licensing concept will emerge, similar to the idea of music synchronization rights -- maybe call it "training rights." It won't matter whether the text was purchased or pirated -- just like it doesn't matter now if an audio track was purchased or pirated, when it's mixed into in a movie soundtrack. Talent agencies will negotiate training rights fees in bulk for popular content creators, wh…

>> Talent agencies will negotiate training rights fees in bulk for popular content creators AFAICT there is no legal recognition of "training rights" or anything similar. First sale right is a thing, but even textbooks don't get extra rights for their training or educational value.

This is why jmkb referenced synchronization rights, which (as I recall) were invented when they seemed useful. jmkb is suggesting a new right might be created, not claiming that they already exist.

(even if it wasn’t sync rights, there was something else musically related that was created in response to technological development. wikipedia will have plenty on it)

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#179

Sarah's pov raises some questions for me regarding my own "training", there is a noteworthy part of who I am built upon the consumed music, books, movies, video games and tv shows that myself or people around me have pirated and shared with me. This part of me helped me in life appreciably, I could also say I profited because of it, helping me along my life in being likable, funny, relatable, with broad outlooks etc.…

[deleted]
Post reply on HN