Live data from Hacker News

Judge said Meta illegally used books to build its AI

wired.com

331–340 of 352 posts

Re: Judge said Meta illegally used books to build its AI

#331

Earlier quoted context omitted.

Why does it have to be verbatim? Seriously, this I don't understand. If I produce a terrible shakycam recording of a film while sitting in a movie theater, it's not a verbatim copy, nor is it even necessarily representative of the original work -- muddied audio, audience sounds, cropped screen, backs of heads -- and yet it would be considered copyright infringement? How many times does one need to compress the JPEG b…

The purpose of copyright is it progress the arts and sciences. Not to guarantee profit. Guaranteeing profit is just the way we encourage people to progress the arts and sciences. That is why so called derivative works are allowed (and even encouraged). If copyrighted material is ingested, modified or enhanced to add value, and then regurgitated that is legal, whereas copying it without adding value is not legal. If d…

> The purpose of copyright is it progress the arts and sciences. Not to guarantee profit.

That seems to go against the notion that copyright can last beyond the author's lifetime - most arts and science progress tends to reduce after death

Re: Judge said Meta illegally used books to build its AI

#332

Earlier quoted context omitted.

No transformation is needed. The point here is that book files have to be copied before they can be used for training. Copyright texts typically say something like "No unauthorised copying or transmission in any form (physical, electronic, etc.)" Individuals who torrented music and video files have been bankrupted for doing exactly this. The same laws should apply when a corporation downloads torrent files. What happ…

> have been bankrupted for doing exactly this. Only if they seeded the data and some other entity downloaded it, i.e. they hosted the data. In a previous article I believe it was called out that Meta was being a leecher (not seeding back what they downloaded). It's the hosting that gets you, not the act of downloading it.

> It's the hosting that gets you, not the act of downloading it.

However, people have been prosecuted for not even hosting a torrent, but merely providing a link to where people can find it.

e.g. https://torrentfreak.com/operator-of-popcorn-time-info-site-...

Re: Judge said Meta illegally used books to build its AI

#333
post #5

The title for this submission is somewhat misleading. The judge didn't make any sort of ruling, this is just reporting on a pretrial hearing. He also doesn't seem convinced as to how relevant downloading books from LibGen is to the case: > At times, it sounded like the case was the authors’ to lose, with [Judge] Chhabria noting that Meta was “destined to fail” if the plaintiffs could prove that Meta’s tools created s…

The RIAA lawyers never had to demonstrate that copying a DVD cratered the sales of their clients. They just got high penalties for infringers almost by default. Now that big capital wants to steal from individuals, big capital wins again. (Unrelatedly, has Boies ever won a high profile lawsuit? I remember him from the Bush/Gore recount issue, where he represented the Democrats.)

when kids during the napster era were downloading music, the mega-corporations yelled that their bottom-lines to shareholders were being destroyed.

now, when the mega-corporations do it, it is 'just the cost of doing business'.

in both cases, the mega-corporations win because...they have the most money. law, and certainly justice, is not for the poor. at least not in america.

Re: Judge said Meta illegally used books to build its AI

#334

Earlier quoted context omitted.

There was a relatively famous Google case regarding their digitization of books without the authors consent in 2015. Although it's not a perfectly analogous to this situation. In Googles case they were digitizing the books (that they did not own), and publishing snippets for search users to help them find books and other material that weren't indexed on the web. The court found they had that right, but did place some…

According to your link and this comment https://news.ycombinator.com/item?id=43899406 Google's scanning project was ruled fair use because it was "transformative" and didn't harm the market for the works. It allowed searching within books that was otherwise impossible at the time, but by not providing the full text of the books it didn't meaningfully reduce sales. Someone photocopying a book to read on the toilet (an…

> Someone photocopying a book to read on the toilet (and leave the original on their nightstand) isn't engaging in transformative use.

The situation you describe is more akin to "format shifting" or maybe "space shifting", which is converting copyrighted material to other formats or places as a backup, for preservation purposes, or just for convenience. It is legally protected in the US, and most of the rest of the world. I believe that was settled due to early litigation with VCR's, but there are mountains of relevant cases at this point.

I do recall reading that the UK recently changed their laws in that regard, and like most copyright law it can depend on the specifics of the situation and even (*yuck*) intent. So it's always worth checking if you're unsure.

Re: Judge said Meta illegally used books to build its AI

#335

Earlier quoted context omitted.

In the general case, yes, but they can verifiably reproduce at least some copyrighted works verbatim, which implies, at the minimum, that their content is stored in model weights in some fashion.

Everyone knows the training data is stored in some way in the LLM. The point is the use of the copyrighted material is transformative. Remember google books, it literally shows photocopy of pages of books but the court ruled it’s fair use. A simplified explanation is book vs search engine and book vs ai chatbot are very different from each other.

"a photocopy of pages of books" is exactly that, pages of an existing book. It doesn't pretend to be something else.

The output of a LLM, when based heavily on that same page, pretends to be something novel. IMadeThis_Meme.gif

Re: Judge said Meta illegally used books to build its AI

#336

Earlier quoted context omitted.

Information contained in subjective mental experience is not a medium which can qualify as a copy (infringing or otherwise) in copyright law, whereas data recorded in digital media such as computer memory is , so they are not similar circumstances with regard to copyright law, however analogous you might feel they are from some other perspective.

What if there was a camera watching the student read in the library, such that the words could be seen?

That doesn't change anything about the student reading.

The recording made by the camera would be a whole different issue.

Re: Judge said Meta illegally used books to build its AI

#337

Earlier quoted context omitted.

Then the results would be the same, and it would still be fair use. I have yet to see an example that demonstrates LLMs plagiarize by default or by tendency. Your causality seems to be inverted here. You seem to be implying that "learning" (or the ingestion and retention of information for the same means) is banned by default for everything, but we decide to allow it for humans as the sole exception. This is not the…

No, it wouldn't. Because if I record "Revenge of the Sith", compress it, and then distribute it for free online, that's obviously not fair use. Fair use is pretty complicated. Part of Fair Use is the "The Effect of the Use on the Potential Market for or Value of the Work", which already puts even human commercial endeavors in a tough spot. You can make it work, but you have to really try. Satire like Weird Al or what…

> The only reason we're even really entertaining this is because people continually draw parallels to humans. You see, it's not stealing from Getty. It's more like if someone saw Getty Images and then went out and took a photo in that same flat, boring style. Except nobody saw anything. And nobody went out an took a photo.

But unless your argument is that the photo outputs from the GenAI are literally equivalent to the training data, you would agree the end result is the same, right? Anyone can see that the images are not the training data stitched together, so it doesn't even really matter how it all works mechanistically, even though your description ("glorified database") is wrong.

Re: Judge said Meta illegally used books to build its AI

#338

Earlier quoted context omitted.

Information contained in subjective mental experience is not a medium which can qualify as a copy (infringing or otherwise) in copyright law, whereas data recorded in digital media such as computer memory is , so they are not similar circumstances with regard to copyright law, however analogous you might feel they are from some other perspective.

> whereas data recorded in digital media such as computer memory is Exactly. Plus, Thomson Reuters recently won a case regarding using copyrighted material for AI training: https://www.wired.com/story/thomson-reuters-ai-copyright-law... > Thomson Reuters has won the first major AI copyright case in the United States. In 2020, the media and technology conglomerate filed an unprecedented AI copyright lawsuit against th…

If I remember right, that particular case was Ross Intelligence almost copying a database of Westlaw's summaries completely and with the intent to compete directly too, with their only real addition being adding AI to the search capabilities, making it much more nuanced of a decision than just ruling against the use of AI.

Re: Judge said Meta illegally used books to build its AI

#339

Earlier quoted context omitted.

No, it wouldn't. Because if I record "Revenge of the Sith", compress it, and then distribute it for free online, that's obviously not fair use. Fair use is pretty complicated. Part of Fair Use is the "The Effect of the Use on the Potential Market for or Value of the Work", which already puts even human commercial endeavors in a tough spot. You can make it work, but you have to really try. Satire like Weird Al or what…

> The only reason we're even really entertaining this is because people continually draw parallels to humans. You see, it's not stealing from Getty. It's more like if someone saw Getty Images and then went out and took a photo in that same flat, boring style. Except nobody saw anything. And nobody went out an took a photo. But unless your argument is that the photo outputs from the GenAI are literally equivalent to t…

Part of my point is that you don't need to produce literally equivalent output. Again, if I record and compress "Revenge of the Sith", there's literally zero pixels shared between my recording and the actual movie. Cool, so I can go upload it for free then right? No, I can't.

Can GenAI produce indistinguishable images to what's on Getty Images? If you write the prompt correctly, yes. I know because there are services where you can get generated stock images.

Re: Judge said Meta illegally used books to build its AI

#340

Earlier quoted context omitted.

> The only reason we're even really entertaining this is because people continually draw parallels to humans. You see, it's not stealing from Getty. It's more like if someone saw Getty Images and then went out and took a photo in that same flat, boring style. Except nobody saw anything. And nobody went out an took a photo. But unless your argument is that the photo outputs from the GenAI are literally equivalent to t…

Part of my point is that you don't need to produce literally equivalent output . Again, if I record and compress "Revenge of the Sith", there's literally zero pixels shared between my recording and the actual movie. Cool, so I can go upload it for free then right? No, I can't. Can GenAI produce indistinguishable images to what's on Getty Images? If you write the prompt correctly, yes. I know because there are service…

> Part of my point is that you don't need to produce literally equivalent output. Again, if I record and compress "Revenge of the Sith", there's literally zero pixels shared between my recording and the actual movie. Cool, so I can go upload it for free then right? No, I can't.

That's because you would be redistributing the actual material, just in a really roundabout way. GenAI models are not that, they're not a database and don't work like one.

> Can GenAI produce indistinguishable images to what's on Getty Images?

That doesn't matter because you can't copyright a style. From the point of view of copyright law, it would look like you were copying nothing proprietary/owned at all.

Post reply on HN