Earlier quoted context omitted.
Movie reviews are fair use because they don't compete with the original work.
The AI models also do not compete with the original work. Nobody is going to try to extract pages by pages a book from ChapGPT, let's be realistic. (and you can't anyways)
Judge said Meta illegally used books to build its AI
341–350 of 352 posts
Re: Judge said Meta illegally used books to build its AI
#342Earlier quoted context omitted.
The AI models also do not compete with the original work. Nobody is going to try to extract pages by pages a book from ChapGPT, let's be realistic. (and you can't anyways)
Half the point of this crap is labor automation. How does it not compete? If I write a bunch of books and then you use my work to make a machine that writes books in my style, you are using my labor to directly compete with me.
I don't think the current state of LLM would be able to write 200 pages in a coherent manner anyways.
Re: Judge said Meta illegally used books to build its AI
#343Let me make a clarifying statement since people confuse (purposely or just out of ignorance) what violating copyright for AI training can refer to: 1. Training AI on freely available copyright - Ambiguous legality, not really tested in court. AI doesn't actually directly copy the material it trains on, so it's not easy to make this ruling. 2. Circumventing payment to obtain copyright material for training - Unambiguo…
If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. What is inspiration? What is imitation? What is plagiarism? The lines aren't clearly drawn for humans... much less for LLMs.
Re: Judge said Meta illegally used books to build its AI
#344Earlier quoted context omitted.
One could argue a student going to a library and reading a book is analogous and legal.
Information contained in subjective mental experience is not a medium which can qualify as a copy (infringing or otherwise) in copyright law, whereas data recorded in digital media such as computer memory is , so they are not similar circumstances with regard to copyright law, however analogous you might feel they are from some other perspective.
Pirating the book is copyright infringement ... reading in a library is not.
Training of a neural network on copyrighted work is the algorithmic equivalent of a "subjective mental experience".
Re: Judge said Meta illegally used books to build its AI
#345Earlier quoted context omitted.
The RIAA lawyers never had to demonstrate that copying a DVD cratered the sales of their clients. They just got high penalties for infringers almost by default. Now that big capital wants to steal from individuals, big capital wins again. (Unrelatedly, has Boies ever won a high profile lawsuit? I remember him from the Bush/Gore recount issue, where he represented the Democrats.)
Meta needs these books. They seek to convert them into more products. The needs of the copyright holders , who are relatively small businesses and individuals are outweighed by the needs of Meta. Sarah wanting to watch a movie or listen to music... Too bad she doesn't have an elite team of lawyers to justify whatever she wants. In practice Meta has the money to stretch this out forever and at most pay inconsequential…
Edit: I didn't make it clear... I don't think meta is going to be paying or offering a revenue stream like YouTube ended up creating. I also have no idea if YT actually brings in money for those groups and if the copyright holders essentially took what they could get or were happy with the deal so who knows.
It's only when other parts of the system get abused there's problems but that's a sep issue...
Re: Judge said Meta illegally used books to build its AI
#346Earlier quoted context omitted.
OK, so we name it something different, you transform inputs into smaller outputs. If I make a script without AI that transforms someones poems without permission so sometimes it outputs the exact poems but sometimes it does it wrong, when is my script fair to use and when what I did is illegal. say my script contains words and matrixes of numbers the original poems are not directly inside, the script transformed them…
YouTube actively filters small channel copied content too. We should be building robust copyright filters and everyone should be able to contribute their work to it. but that is a different issue than whether or not an LLM is legally allowed to view a work that is publicly available. Again, pretty much every artist is capable of off-hand copyright violation on the spot. This has been true forever. We don't bar them f…
We are not talking about the scenario where LLM does a web search and "transforms" your blog as a response to a user prompt.
We are talking about people writing code to torrent content, including copyrighted content as books and newspaper articles then this same people create soem software/script/blackbox and they put this copyrighted stuff inside and this same people also put a filter in the output to attempt and block SOME of the copyrighted material to be spitted out(they did it for Dune, Harry Poter , probably other popular stuff but for sure they did not done it for Romanian copyrighted material or some small writer).
If I publish a poem I do not give implicit permission to anyone to use it as they see fit, and have scripts/software/backboxes train on it and spit parts of my poem out.
Re: Judge said Meta illegally used books to build its AI
#347Earlier quoted context omitted.
According to your link and this comment https://news.ycombinator.com/item?id=43899406 Google's scanning project was ruled fair use because it was "transformative" and didn't harm the market for the works. It allowed searching within books that was otherwise impossible at the time, but by not providing the full text of the books it didn't meaningfully reduce sales. Someone photocopying a book to read on the toilet (an…
> Someone photocopying a book to read on the toilet (and leave the original on their nightstand) isn't engaging in transformative use. The situation you describe is more akin to "format shifting" or maybe "space shifting", which is converting copyrighted material to other formats or places as a backup, for preservation purposes, or just for convenience. It is legally protected in the US, and most of the rest of the w…
Citation needed.
Re: Judge said Meta illegally used books to build its AI
#348Earlier quoted context omitted.
No, because the entire argument hinges on the fact that LLMs learn, which is like humans learning, so it's transformative. That only works if you consider learning or transformation to be something that does not rely on the human spirit. Which, actually, most people do not believe. And it's pretty difficult to argue - we don't even know how learning works for people. A lot of people just jump to LLMs learning like it…
> That only works if you consider learning or transformation to be something that does not rely on the human spirit. Even changes made using simple non-ML algorithms can be transformative according to fair use doctrine, like the thumbnailing of images done by search engines. It's not meant in some spiritual sense.
But, a lot of AI products are specifically and explicitly designed to obsolesce the thing they trained off. No need to go to Encyclopedia X or the NYT, this has the same content.
Re: Judge said Meta illegally used books to build its AI
#349Earlier quoted context omitted.
Part of my point is that you don't need to produce literally equivalent output . Again, if I record and compress "Revenge of the Sith", there's literally zero pixels shared between my recording and the actual movie. Cool, so I can go upload it for free then right? No, I can't. Can GenAI produce indistinguishable images to what's on Getty Images? If you write the prompt correctly, yes. I know because there are service…
> Part of my point is that you don't need to produce literally equivalent output. Again, if I record and compress "Revenge of the Sith", there's literally zero pixels shared between my recording and the actual movie. Cool, so I can go upload it for free then right? No, I can't. That's because you would be redistributing the actual material, just in a really roundabout way. GenAI models are not that, they're not a dat…
Right, which I’m arguing is what LLMs do just in an even more roundabout way.
The technical details of LLMs don’t actually matter. We don’t really care if they’re a database or not. The question is do they reproduce the source material? And yeah, pretty much they do, in a lot of instances. Not all, but a lot.
To produce yet another analogy, imagine I have a service X. You can pay and I will give you any movie you want. You don’t know how I do it. Is this copyright infringement or not? I would say yes. Now let’s say I reveal the secret - I open up photoshop and painstakingly recreate the movie frame by frame. I might make a mistake here or there. Is this still copyright infringement? I think it is.
Re: Judge said Meta illegally used books to build its AI
#350Earlier quoted context omitted.
> That only works if you consider learning or transformation to be something that does not rely on the human spirit. Even changes made using simple non-ML algorithms can be transformative according to fair use doctrine, like the thumbnailing of images done by search engines. It's not meant in some spiritual sense.
The reason that’s okay is because you aren’t competing against the initial source material. A thumbnail on Google for “Revenge of the Sith” is not a replacement for watching the movie. But, a lot of AI products are specifically and explicitly designed to obsolesce the thing they trained off. No need to go to Encyclopedia X or the NYT, this has the same content.
I don't mean to claim that search engine image thumbnailing is like-to-like in every consideration, just that it demonstrates there's no "human spirit" required in order to qualify as "transformative" as far as fair use is concerned. Search engine image thumbnailing has been found to be transformative, for instance in Perfect 10, Inc. v. Amazon.com, Inc.: "Google's use of thumbnails is highly transformative."
And, though I'm probably being pedantic here, I think it's important to distinguish that the other fair use factor you allude to is not whether you're "competing against" the original work, but specifically the effect of your use on the market/value of that original work. For example if your documentary uses a clip from a TV show and also happens to air in the same time-slot as that TV show - the extent you compete/displace market for the TV show in general (even as you would had you not included the clip of it) is not what's under consideration, but rather only the additional extent you displace its market specifically due to inclusion of that clip.
Because of that, I'd claim that some machine-learning-based tool that partially displaces the market for a work it was trained on (for instance, Google Translate displacing the market for a translated version of a book) might still be seen reasonably favorably under the market impact factor, so long as the extent it displaces that work is largely independent of whether it has trained on that work specifically (such as if the translation tool could already provide a decent translation of the original book even before having trained on its translated version).