Live data from Hacker News

Judge said Meta illegally used books to build its AI

wired.com

81–90 of 352 posts

Re: Judge said Meta illegally used books to build its AI

#81

Earlier quoted context omitted.

> AI doesn't actually directly copy the material it trains on Of course it does. Large models are trained on gigantic clusters. How can you train without copying the material to machines in the cluster?

“Copy” is ambiguous here. Of course data is copied during training. That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data.

Why does it have to be verbatim? Seriously, this I don't understand.

If I produce a terrible shakycam recording of a film while sitting in a movie theater, it's not a verbatim copy, nor is it even necessarily representative of the original work -- muddied audio, audience sounds, cropped screen, backs of heads -- and yet it would be considered copyright infringement?

How many times does one need to compress the JPEG before it's fair use? I'm legitimately curious what the test is here.

Re: Judge said Meta illegally used books to build its AI

#82

Earlier quoted context omitted.

“Copy” is ambiguous here. Of course data is copied during training. That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data.

Copyright is the right to make copies. Why is copying during training is any different from producing copies of training data after training? If we're going that way, let me torrent every movie and TV show ever to "train" myself.

I don't think this is a reasonable argument. I don't think copyright is actually defined in that sense, but is perhaps more focused on consuming the content. Is an http proxy making a copy of something? What about computing an md5 of it as it's streamed through the proxy? Or maybe counting the words in the thing being served in order to track stats? I'd argue none of these fall under copyright, but each is an incremental step towards what it means to train a model.

Re: Judge said Meta illegally used books to build its AI

#83

Earlier quoted context omitted.

“Copy” is ambiguous here. Of course data is copied during training. That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data.

Copyright is the right to make copies. Why is copying during training is any different from producing copies of training data after training? If we're going that way, let me torrent every movie and TV show ever to "train" myself.

It's almost like information wants to be free

Re: Judge said Meta illegally used books to build its AI

#84

Earlier quoted context omitted.

“Copy” is ambiguous here. Of course data is copied during training. That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data.

Why does it have to be verbatim? Seriously, this I don't understand. If I produce a terrible shakycam recording of a film while sitting in a movie theater, it's not a verbatim copy, nor is it even necessarily representative of the original work -- muddied audio, audience sounds, cropped screen, backs of heads -- and yet it would be considered copyright infringement? How many times does one need to compress the JPEG b…

If you read a book and later understand its plot but can only explain it in your own words, did you copy it?

The model isn’t storing the book.

Re: Judge said Meta illegally used books to build its AI

#85

Earlier quoted context omitted.

> Meta did pirate the works but may be entitled to use them under fair use What fair use? Were the books promised to them by god or something?

"fair use" is a specific legal term

In a specific legal jurisdiction.

The Berne convention mentions "fair practice", and puts the responsibility on the individual countries.

Re: Judge said Meta illegally used books to build its AI

#86
post #55

Let me make a clarifying statement since people confuse (purposely or just out of ignorance) what violating copyright for AI training can refer to: 1. Training AI on freely available copyright - Ambiguous legality, not really tested in court. AI doesn't actually directly copy the material it trains on, so it's not easy to make this ruling. 2. Circumventing payment to obtain copyright material for training - Unambiguo…

I'm not sure if Meta did anything illegal in 2. either. I thought the copyright infringement was by the people who provided the copyrighted material when they did not have the rights to do so. I may be wrong on this, but it would seem a reasonable protection for consumers in general. Meta is hardly an average consumer, but I doubt that matters in the case of the law. Having grounds to suspect that the provider did no…

> but it would seem a reasonable protection for consumers in general.

The final say may ultimately come from the Cox vs Record Labels case from 2019 that is still working it's way through the appeal courts.

If the record labels win their appeal, anyone who helped facilitate the infringement can be brought into a lawsuit. The record labels sued Cox for infringement by it's users. It's not out of the question that any ISP that provides Internet connectivity to Facebook could be pulled in for damages.

For Meta these two cases could result in an existential threat to the company, and rightly so because the record labels do not play games. The blood is already in the water.

Re: Judge said Meta illegally used books to build its AI

#87

Earlier quoted context omitted.

"fair use" is a specific legal term

I'm aware. I was unsure what doctrine of fair use meta's behaviour could be defended as. What am I, if not an LLM, ingesting copyrighted materials so that I may improve my own future outputs? Why is my own piracy not protected in the same manner?

> Why is my own piracy not protected in the same manner?

You aren't a multi-billion dollar company

Re: Judge said Meta illegally used books to build its AI

#88

Earlier quoted context omitted.

“Copy” is ambiguous here. Of course data is copied during training. That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data.

Copyright is the right to make copies. Why is copying during training is any different from producing copies of training data after training? If we're going that way, let me torrent every movie and TV show ever to "train" myself.

It depends on your license? I mean strictly speaking if you stream a video you purchase legally over say amazon prime, there's lots of "copying" happening at various levels after those bits leave the data center.

Re: Judge said Meta illegally used books to build its AI

#89

Earlier quoted context omitted.

“Copy” is ambiguous here. Of course data is copied during training. That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data.

Copyright is the right to make copies. Why is copying during training is any different from producing copies of training data after training? If we're going that way, let me torrent every movie and TV show ever to "train" myself.

You can train yourself with every book at any library. You can also train yourself on a large number of movies and TV shows for a small monthly fee.

Where your analogy goes wrong is you're saying you want to "[Circumvent] payment to obtain copyright material for training" to use Workaccount2's words.

Re: Judge said Meta illegally used books to build its AI

#90

Earlier quoted context omitted.

> AI doesn't actually directly copy the material it trains on Of course it does. Large models are trained on gigantic clusters. How can you train without copying the material to machines in the cluster?

“Copy” is ambiguous here. Of course data is copied during training. That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data.

"Of course data is copied during training" is copying. As far as I know, the law is consistent that temporary copies are also covered by the copyright act, and that's how some analogous cases were resolved.
Post reply on HN