Live data from Hacker News

Judge said Meta illegally used books to build its AI

wired.com

251–260 of 352 posts

Re: Judge said Meta illegally used books to build its AI

#251
post #61

Earlier quoted context omitted.

counterpoint: Gorillaz is a band specifically designed around the idea that the artist doesn't have to exist in order to do all of the things you mention above. Gorillaz has an image, a style, a mythos, all of that. Granted, when they started at the time (2000) there needed to be human creativity in order to create all of that but with AI everything about this can now be generated. I suspect it won't be long before a…

Very relevant point. My own counter-example of Hatsune Miku is not totally AI generated - just generated using very sophisticated musical tools (see: Vocaloid). There's some very impressive youtubers who are claiming to be generating new music with AI. The one I listen to the most I very much doubt has everything 100% generated - he probably generates a bunch of melodies and other bits of track and stiches the best c…

Every "Hatsune Miku" song has a talented person behind that avatar. Miku is just a synth and 3D model.

Re: Judge said Meta illegally used books to build its AI

#252

Earlier quoted context omitted.

>That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data. The NYTimes in 2023 was able to demonstrate that the models can reproduce entire articles verbatim[0] with minimal coercion. [0] https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20...

Perhaps this is not evidence that the NY Times article was copied, but that what the NY TImes writes is highly predictable.

The obvious test for this would be to have the models produce an article from before and after its cutoff date and see if the output is still verbatim.

It would be a remarkable quirk of statistics that, if given all text on the internet except for the NYTimes back catalogue, a model would produce any NYT article.

Re: Judge said Meta illegally used books to build its AI

#253

Earlier quoted context omitted.

> And it's nearly impossible to play a CD without that copying. Exactly. It's how you are supposed to use the CD. That's not true for your example of a book on a webcam. You're supposed to read the book, not an image of the book. That's not the same thing as copying for "processing".

Well you said explicit license earlier. Would it be a violation to play back a record like a CD and have a digital buffer? That would be pretty silly.

> explicit license

Selling you the CD explicitly grants you the right to play it using a CD player.

> Would it be a violation to play back a record like a CD and have a digital buffer?

I genuinely have no idea lol.

Re: Judge said Meta illegally used books to build its AI

#254

Earlier quoted context omitted.

So Meta borrowed every book from a library and paid to obtain all of the movies and TV shows? They kept only one copy of every book at any time on their system? Because I'm certainly not allowed to photocopy a library book in its entirety. And I guarantee you a Netflix subscription doesn't allow me to keep a copy of a movie on my hard drive and use it for training man or machine.

> Because I'm certainly not allowed to photocopy a library book in its entirety. IANAL but that probably falls under fair use? You'll get in trouble if you photocopy the work and sell access to it.

> but that probably falls under fair use?

I've not found case law for that. I've had this same argument on HN multiple times over the past few months.

Re: Judge said Meta illegally used books to build its AI

#255
post #77

Earlier quoted context omitted.

If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. What is inspiration? What is imitation? What is plagiarism? The lines aren't clearly drawn for humans... much less for LLMs.

> If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. I can absolutely guarantee you that neither DeepSeek nor Alibaba's highly talented Qwen group will care even a little bit, in the long run. Not if there's value to be had in AI. (And I can tell you down to the dollar what LLMs can save in certain business use cases.) If the…

> If the US decides to unilaterally shut down LLMs, that just means that the rest of the world will route around us.

You're talking as if they are some kind of nationalized or publically-owned asset, as opposed to a bunch of for-profit, privately-owned silos.

Re: Judge said Meta illegally used books to build its AI

#256
post #255
post #77

Earlier quoted context omitted.

> If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. I can absolutely guarantee you that neither DeepSeek nor Alibaba's highly talented Qwen group will care even a little bit, in the long run. Not if there's value to be had in AI. (And I can tell you down to the dollar what LLMs can save in certain business use cases.) If the…

> If the US decides to unilaterally shut down LLMs, that just means that the rest of the world will route around us. You're talking as if they are some kind of nationalized or publically-owned asset, as opposed to a bunch of for-profit, privately-owned silos.

Local models are a thing, though. You can run DeepSeek on your local computer.

Even if ChatGPT, Huggingface, etc. died, we would still have the models and we would still be able to run them.

Re: Judge said Meta illegally used books to build its AI

#257

Earlier quoted context omitted.

They are in no way compression algorithms. They can be modeled like that in the same way you can model humans as lossy compression algorithms. You would never use a human to backup your financial reports, but the human might be able to give a good overview. You would never use an LLM to backup your financial reports, but they might be able to give a good overview. AI training data is disposable. There is nothing that…

>It's 20PB of examples used to form the shape of a 20GB model. You could show it 5GB of training data or 500EB of training data and it would still be 20GB - because it is not a compression algo, it's a 20GB shape formed by external data. You can compress 20PB of text to 20Gb or even less, if input is super repetitive. So the same with images, if 50% of the images are cats then you learn how to represent the cat pixel…

The point is that it isn't compression. Its molding a plain structure iteratively into a ultra complex one. The model starts and ends at 20GB. It might have features that are reminiscent of compression or act like it, but under the hood there is nothing like zip, rar, H.265, or JPEG going on.

And yes LLMs can recall exact material, but it is excerpts and fragments. There is statistical significance to it's ordering. Humans readily do this too (excerpts and fragments), most artists can draw a batman symbol (but not an episode of batman). That doesn't in anyway mean that artists should not be allowed to ever see a batman symbol. It means that artists shouldn't be allowed to get paid to draw one. And they are not. And LLMs are not exempt either.

But the fix is output filtering, just like everything else that can violate copyright. Which is already being done (albeit poorly, but way better than 2 years ago), the same as artists will not draw the batman symbol for you despite being able to.

Re: Judge said Meta illegally used books to build its AI

#258
post #77

Earlier quoted context omitted.

If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. What is inspiration? What is imitation? What is plagiarism? The lines aren't clearly drawn for humans... much less for LLMs.

> If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. I can absolutely guarantee you that neither DeepSeek nor Alibaba's highly talented Qwen group will care even a little bit, in the long run. Not if there's value to be had in AI. (And I can tell you down to the dollar what LLMs can save in certain business use cases.) If the…

Or AI companies could use some of their vast reserves of cash to pay for licensing agreements and pay people for their fucking intellectual property, then feed it to the beast.

But then they'd have to actually communicate with people and negotiate consent instead of just hoovering up everything they can get their hands on in their quest to replace it.

Re: Judge said Meta illegally used books to build its AI

#259

Earlier quoted context omitted.

> They are in no way compression algorithms. I'm sorry, but this a fundamentally incorrect view of machine learning (including, but not limited to transformers). From an information theoretic perspective the two are essentially identical with the exception that standard compression algorithms do not have a proper "loss" function other than just trying to minimize reconstruction loss with the resulting compression siz…

>They can be modeled like that in the same way you can model humans as lossy compression algorithms Humans are totally capable of data compression. This will just devolved into a semantics game of what a data compressor is. LLMs were not developed to be, do not function as, and are not use as data compression utilities. Please, come knocking when a service provider exists that will use LLM's to compactly store your c…

> LLMs were not developed to be, do not function as, and are not use as data compression utilities.

Again, from a information theoretic view point, this is exactly what they are doing, how they where developed and how they function.

I don't know any serious researcher in ML that would find this claim even remotely controversial. It's really not just "a semantics game", its a part of a foundational understanding of the topic. If you want to understand LLMs from this perspective, a good place to start is with an auto-encoder which does try to learn a standard compression algorithm, the move on to more sophisticated embedding models (found in a lot of recommender systems) which try to learn an additional objective on top of minimizing reconstruction error. You'll then see that Transformers and all other major NN architectures fall out of these basic principles.

> Please, come knocking when a service provider exists that will use LLM's to compactly store your company data.

This is literally what every vectordb company does right now, as well as all "chat with your docs" type startups.

Re: Judge said Meta illegally used books to build its AI

#260

Earlier quoted context omitted.

>It's 20PB of examples used to form the shape of a 20GB model. You could show it 5GB of training data or 500EB of training data and it would still be 20GB - because it is not a compression algo, it's a 20GB shape formed by external data. You can compress 20PB of text to 20Gb or even less, if input is super repetitive. So the same with images, if 50% of the images are cats then you learn how to represent the cat pixel…

The point is that it isn't compression. Its molding a plain structure iteratively into a ultra complex one. The model starts and ends at 20GB. It might have features that are reminiscent of compression or act like it, but under the hood there is nothing like zip, rar, H.265, or JPEG going on. And yes LLMs can recall exact material, but it is excerpts and fragments. There is statistical significance to it's ordering.…

OK, so we name it something different, you transform inputs into smaller outputs. If I make a script without AI that transforms someones poems without permission so sometimes it outputs the exact poems but sometimes it does it wrong, when is my script fair to use and when what I did is illegal. say my script contains words and matrixes of numbers the original poems are not directly inside, the script transformed them into vectors.

maybe even simpler, I create a zip format where I randomly replace words with their synonym, or group of words with something equivalent. Would you defend this as original ? why my random transformations are not original while mathematical transformations you will defend ?

And how can you suggest putting output filters to protect only the giants for copyright and everyone else gets screwed.

Post reply on HN