Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

21–30 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#25

Isn’t it much more likely that there are a lot of book reviews and summaries in its training set from which it can synthesize its own?

I think it's actually much more likely that they just dumped a bunch of book PDFs in the training folder and let it go to work. I seriously doubt any of these AI companies are being even the least bit careful about the data they're lapping up for training

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#26

Isn’t it much more likely that there are a lot of book reviews and summaries in its training set from which it can synthesize its own?

Where are the summaries from? I would say it's much more likely that a shadow library was scraped but if course that is also seemingly impossible to prove. One may be able to somewhat test that by asking for a summary of a book/text only available on a shadow library.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#27

Earlier quoted context omitted.

I'm allowed to make private copies of copywritten works. I'm not allowed to redistribute them. To what extent this is redistribution is not clear. Is there much of difference between this model and a machine, like a VCR, that recreates the original work when I press a button?

This would be like you intensely studying the copy written work and then writing things based on the knowledge you obtained from that. Except, we don't know if their is an exception for things learned by people vs. things learned by machines, or if the machines are not really learning but copying instead (or if learning is intrinsically a form of copying?).

In the case of unreleased work, you writing about your knowledge of it is just proof that you obtained the work, which is proof that you committed a tort/trespass. Just like if you published a newspaper article with information you could only have acquired by hacking someone's phone. I'm not sure what a court would find against you, but it seems clear that there would be some way to couch that as a legal grievance.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#28

Isn’t it much more likely that there are a lot of book reviews and summaries in its training set from which it can synthesize its own?

Except they have a documented paper trail showing illegal book repos were used in training

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#30

This is actually quite interesting, as it's drawing a distinction between training material that can be accessed by anybody with a web browser (like anybody's blog), vs. training material that was "illegally-acquired... available in bulk via torrent systems." I don't think there's any reason why this would be a relevant legal distinction in terms of distributing an LLM -- blog authors weren't giving consent either. H…

I'm allowed to make private copies of copywritten works. I'm not allowed to redistribute them. To what extent this is redistribution is not clear. Is there much of difference between this model and a machine, like a VCR, that recreates the original work when I press a button?

Is Game of Thrones a redistribution of Lord of the Rings?
Post reply on HN