Live data from Hacker News

Judge said Meta illegally used books to build its AI

wired.com

121–130 of 352 posts

Re: Judge said Meta illegally used books to build its AI

#121
post #90

Earlier quoted context omitted.

“Copy” is ambiguous here. Of course data is copied during training. That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data.

"Of course data is copied during training" is copying. As far as I know, the law is consistent that temporary copies are also covered by the copyright act, and that's how some analogous cases were resolved.

If I buy a book I'm free to print as many copies as I want inside my house

It becomes illegal if I try to distribute those copies

So the question is, does distributing an AI that has been trained on Harry Potter count as distributing Harry Potter?

Re: Judge said Meta illegally used books to build its AI

#122

Earlier quoted context omitted.

Why does it have to be verbatim? Seriously, this I don't understand. If I produce a terrible shakycam recording of a film while sitting in a movie theater, it's not a verbatim copy, nor is it even necessarily representative of the original work -- muddied audio, audience sounds, cropped screen, backs of heads -- and yet it would be considered copyright infringement? How many times does one need to compress the JPEG b…

There’s something called a substantive transformation test in copyright law. When you write a summary of a book, you don’t infringe on copyright because it’s a “substantial transformation”. This goes along with the idea that you can copyright the text but not the ideas it expresses. When model training reads the text and creates weights internally, is that a substantial transformation? I think there’s a pretty strong…

No transformation is needed.

The point here is that book files have to be copied before they can be used for training. Copyright texts typically say something like "No unauthorised copying or transmission in any form (physical, electronic, etc.)"

Individuals who torrented music and video files have been bankrupted for doing exactly this.

The same laws should apply when a corporation downloads torrent files. What happens to them after they're downloaded is irrelevant to the argument.

If this is enforced (still to be seen...) it would be financially catastrophic for Meta, because there are set damages for works that have been registered for copyright protection - which most trad-pubbed books, and many self-pubbed books, are.

Re: Judge said Meta illegally used books to build its AI

#123
post #90

Earlier quoted context omitted.

“Copy” is ambiguous here. Of course data is copied during training. That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data.

"Of course data is copied during training" is copying. As far as I know, the law is consistent that temporary copies are also covered by the copyright act, and that's how some analogous cases were resolved.

Temporary copies are in the scope of copyright law, yes. But also, you are allowed to make them. Or reading a book via a computer would be illegal.

Re: Judge said Meta illegally used books to build its AI

#124

Earlier quoted context omitted.

The output of the LLM is very different from the original, though. It’s hard to look at this and claim it isn’t.

> The output of the LLM is very different from the original, though I will disagree with that characterisation. IMHO: In some cases no, it's not different, there are clear lines from inputs to output. In some cases yes, it's different from any one input work, it's distributed micro-plagiarism of a huge number of sources. In no case is it original. But I think that this is legally undecided and won't be decided by you…

> In some cases yes, it's different from any one input work, it's distributed micro-plagiarism of a huge number of sources. In no case is it original.

That’s like saying the dictionary is micro-plagiarism of a huge number of sources because it uses all the words from those sources.

Re: Judge said Meta illegally used books to build its AI

#125

Earlier quoted context omitted.

“Copy” is ambiguous here. Of course data is copied during training. That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data.

>That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data. The NYTimes in 2023 was able to demonstrate that the models can reproduce entire articles verbatim[0] with minimal coercion. [0] https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20...

Perhaps this is not evidence that the NY Times article was copied, but that what the NY TImes writes is highly predictable.

Re: Judge said Meta illegally used books to build its AI

#126

Earlier quoted context omitted.

Why does it have to be verbatim? Seriously, this I don't understand. If I produce a terrible shakycam recording of a film while sitting in a movie theater, it's not a verbatim copy, nor is it even necessarily representative of the original work -- muddied audio, audience sounds, cropped screen, backs of heads -- and yet it would be considered copyright infringement? How many times does one need to compress the JPEG b…

There’s something called a substantive transformation test in copyright law. When you write a summary of a book, you don’t infringe on copyright because it’s a “substantial transformation”. This goes along with the idea that you can copyright the text but not the ideas it expresses. When model training reads the text and creates weights internally, is that a substantial transformation? I think there’s a pretty strong…

This is a leap in the argument. We've gone from the right to use a work to "unless the result is identical or close to it, we have full rights to all works.".

Seems like a big gap there.

Re: Judge said Meta illegally used books to build its AI

#127

Earlier quoted context omitted.

> The output of the LLM is very different from the original, though I will disagree with that characterisation. IMHO: In some cases no, it's not different, there are clear lines from inputs to output. In some cases yes, it's different from any one input work, it's distributed micro-plagiarism of a huge number of sources. In no case is it original. But I think that this is legally undecided and won't be decided by you…

> In some cases yes, it's different from any one input work, it's distributed micro-plagiarism of a huge number of sources. In no case is it original. That’s like saying the dictionary is micro-plagiarism of a huge number of sources because it uses all the words from those sources.

I disagree, you can't ask a dictionary to "generate 2000 words in the style of (author)".

Re: Judge said Meta illegally used books to build its AI

#128
post #5

The title for this submission is somewhat misleading. The judge didn't make any sort of ruling, this is just reporting on a pretrial hearing. He also doesn't seem convinced as to how relevant downloading books from LibGen is to the case: > At times, it sounded like the case was the authors’ to lose, with [Judge] Chhabria noting that Meta was “destined to fail” if the plaintiffs could prove that Meta’s tools created s…

The RIAA lawyers never had to demonstrate that copying a DVD cratered the sales of their clients. They just got high penalties for infringers almost by default. Now that big capital wants to steal from individuals, big capital wins again. (Unrelatedly, has Boies ever won a high profile lawsuit? I remember him from the Bush/Gore recount issue, where he represented the Democrats.)

That's because that was statutory infringement where marketplace impact things come up more in fair use ("drummer reacts to hearing most famous drummer for the first time"). They look at whether it acts as a substitute for the original, but there are different rules depending on the type of fair use, how transformative it is, and more.

Re: Judge said Meta illegally used books to build its AI

#129

Let me make a clarifying statement since people confuse (purposely or just out of ignorance) what violating copyright for AI training can refer to: 1. Training AI on freely available copyright - Ambiguous legality, not really tested in court. AI doesn't actually directly copy the material it trains on, so it's not easy to make this ruling. 2. Circumventing payment to obtain copyright material for training - Unambiguo…

If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. What is inspiration? What is imitation? What is plagiarism? The lines aren't clearly drawn for humans... much less for LLMs.

> If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another.

You assume that getting tested means the AI trainers lose, and also thar the model architectures that have been developed can’t be retrained from scratch with public domain, owned, and purpose-licensed material. (With several AI companies having been actively pursuing deals to license content for AI training for a while now.)

Re: Judge said Meta illegally used books to build its AI

#130

Earlier quoted context omitted.

Copyright is defined in law and as the original poster stated, whether this is 'copying' as defined by copyright law is legally ambiguous. Copyright doesn't protect against all forms of duplication. For instance, you own the copyright to your post and grant HN a license to offer copies of it. I have no direct license from you to copy the content of your post; but I can copy it to memory, copy a cache to disk, and mak…

> For instance, you own the copyright to your post and grant HN a license to offer copies of it. It’s not a good example, because if you grant a license you give them the right to make copies. The problem is not when Meta got licenses, it’s when they did not.

This line of conversation is not specific to the pirated books but is making claims about AI training in general.
Post reply on HN