Live data from Hacker News

Judge said Meta illegally used books to build its AI

wired.com

141–150 of 352 posts

Re: Judge said Meta illegally used books to build its AI

#141
post #12

> “What about the next Taylor Swift?” he asked, arguing that a “relatively unknown artist” whose work was ingested by Meta would likely have their career hampered if the model produced “a billion pop songs” in their style. I have this debate with a friend of mine. He's terrified of AI making all of our jobs obsolete. He's a brilliant musician and programmer both, so he's both enthused and scared. So let's go with the…

Would Taylor Swift be famous without the support of her label? I sincerely doubt it. How many equally talented artists

We should be careful not to conflate the affects of copyright to the affects of advertising.

Re: Judge said Meta illegally used books to build its AI

#142

Earlier quoted context omitted.

> In some cases yes, it's different from any one input work, it's distributed micro-plagiarism of a huge number of sources. In no case is it original. That’s like saying the dictionary is micro-plagiarism of a huge number of sources because it uses all the words from those sources.

I disagree, you can't ask a dictionary to "generate 2000 words in the style of (author)".

So? Why is that, of all things, the crux of whether it’s copyright infringement?

Plagiarism isn’t necessarily copyright infringement, and plagiarism isn’t illegal. Copyright infringement is.

Even still, your argument that everyone who generates 2,000 words in the style of (author) is plagiarizing is also flatly false. By that standard all English essays that mimic someone else’s style would be plagiarism.

Re: Judge said Meta illegally used books to build its AI

#143

Earlier quoted context omitted.

> AI doesn't actually directly copy the material it trains on Of course it does. Large models are trained on gigantic clusters. How can you train without copying the material to machines in the cluster?

“Copy” is ambiguous here. Of course data is copied during training. That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data.

So if they could produce verbatim segments, that would be a violation? The technology is certainly there and these companies need to work backwards to prevent that.

Re: Judge said Meta illegally used books to build its AI

#145

Earlier quoted context omitted.

Temporary copies are in the scope of copyright law, yes. But also, you are allowed to make them. Or reading a book via a computer would be illegal.

> But also, you are allowed to make them. Not of physical media. You're allowed to make archival copies of digital media. > Or reading a book via a computer would be illegal No you purchased a license (or your library did, in the case of e-borrowing) to read the book on a computer. That makes it legal.

I am allowed to point a webcam at my physical book and read off the screen, even though that makes digital copies of all the text.

Re: Judge said Meta illegally used books to build its AI

#146

Let me make a clarifying statement since people confuse (purposely or just out of ignorance) what violating copyright for AI training can refer to: 1. Training AI on freely available copyright - Ambiguous legality, not really tested in court. AI doesn't actually directly copy the material it trains on, so it's not easy to make this ruling. 2. Circumventing payment to obtain copyright material for training - Unambiguo…

> Circumventing payment to obtain copyright material for training - Unambiguously illegal. The judge in this case seems to disagree with you, not accepting the premise that downloading the material from pirate sites for this use inherently gets the plaintiffs an out from having to address fair use defense as to the actual use. > the plaintiffs want to also tie in the former. No, the defense wants to and the judge has…

> The judge in this case seems to disagree with you, not accepting the premise that downloading the material from pirate sites for this use inherently gets the plaintiffs an out from having to address fair use defense as to the actual use.

This is a good point, as a reminder, the Folsom tests (failing or passing any one is not conclusive, they are to be holistically considered) are:

- the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes (Note also that whether or not the use is transformative is part of this test).

- the nature of the copyrighted work

- the amount and substantiality of the portion used in relation to the copyrighted work as a whole

- the effect of the use upon the potential market for or value of the copyrighted work

https://en.wikipedia.org/wiki/Fair_use#U.S._fair_use_factors

Re: Judge said Meta illegally used books to build its AI

#147
post #55

Let me make a clarifying statement since people confuse (purposely or just out of ignorance) what violating copyright for AI training can refer to: 1. Training AI on freely available copyright - Ambiguous legality, not really tested in court. AI doesn't actually directly copy the material it trains on, so it's not easy to make this ruling. 2. Circumventing payment to obtain copyright material for training - Unambiguo…

I'm not sure if Meta did anything illegal in 2. either. I thought the copyright infringement was by the people who provided the copyrighted material when they did not have the rights to do so. I may be wrong on this, but it would seem a reasonable protection for consumers in general. Meta is hardly an average consumer, but I doubt that matters in the case of the law. Having grounds to suspect that the provider did no…

So you're saying I can legally download movies as long as I don't provide them to others? Sweet!

Re: Judge said Meta illegally used books to build its AI

#148
post #77

Earlier quoted context omitted.

If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. What is inspiration? What is imitation? What is plagiarism? The lines aren't clearly drawn for humans... much less for LLMs.

> If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. I can absolutely guarantee you that neither DeepSeek nor Alibaba's highly talented Qwen group will care even a little bit, in the long run. Not if there's value to be had in AI. (And I can tell you down to the dollar what LLMs can save in certain business use cases.) If the…

The pattern hasn't changed in decades. Remember when ZTE copied Cisco's router code so precisely they included the same bugs and documentation typos?

LLMs are a drop on a hot stone compared to countless other factors why the world already is routing around the US - but I don't want to get political or economical.

Re: Judge said Meta illegally used books to build its AI

#149
post #82

Earlier quoted context omitted.

Copyright is the right to make copies. Why is copying during training is any different from producing copies of training data after training? If we're going that way, let me torrent every movie and TV show ever to "train" myself.

I don't think this is a reasonable argument. I don't think copyright is actually defined in that sense, but is perhaps more focused on consuming the content. Is an http proxy making a copy of something? What about computing an md5 of it as it's streamed through the proxy? Or maybe counting the words in the thing being served in order to track stats? I'd argue none of these fall under copyright, but each is an increme…

> I don't think copyright is actually defined in that sense, but is perhaps more focused on consuming the content.

https://en.wikipedia.org/wiki/American_Broadcasting_Cos.,_In....

I'm not a legal expert. My layman's understanding of the case above is Aereo was in violation because they made copies of content - content that the receiver was already allowed to access - available over the Internet to the intended receiver. That is to say, the copying was the problem.

Re: Judge said Meta illegally used books to build its AI

#150

Let me make a clarifying statement since people confuse (purposely or just out of ignorance) what violating copyright for AI training can refer to: 1. Training AI on freely available copyright - Ambiguous legality, not really tested in court. AI doesn't actually directly copy the material it trains on, so it's not easy to make this ruling. 2. Circumventing payment to obtain copyright material for training - Unambiguo…

If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. What is inspiration? What is imitation? What is plagiarism? The lines aren't clearly drawn for humans... much less for LLMs.

If corporations owned human slaves and fed them copyrighted materials so that they were inspired to produce original creative output, I don't think that creative output should enjoy legal protections either. Even if slavery were not illegal.

Because the obvious question would be - how can free people compete with that?

Post reply on HN