Live data from Hacker News

Judge said Meta illegally used books to build its AI

wired.com

291–300 of 352 posts

Re: Judge said Meta illegally used books to build its AI

#291
post #82

Earlier quoted context omitted.

I don't think this is a reasonable argument. I don't think copyright is actually defined in that sense, but is perhaps more focused on consuming the content. Is an http proxy making a copy of something? What about computing an md5 of it as it's streamed through the proxy? Or maybe counting the words in the thing being served in order to track stats? I'd argue none of these fall under copyright, but each is an increme…

> I don't think copyright is actually defined in that sense, but is perhaps more focused on consuming the content. https://en.wikipedia.org/wiki/American_Broadcasting_Cos.,_In... . I'm not a legal expert. My layman's understanding of the case above is Aereo was in violation because they made copies of content - content that the receiver was already allowed to access - available over the Internet to the intended recei…

I'm not a lawyer either, but from the summary a key part of the case seems to be that they were distributing the video to people for them to watch.

"Aereo's retransmission of television broadcasts was a "public performance" of the networks' copyrighted work. The Copyright Act of 1976 forbids such performances without the permission of the holder of the copyright. Second Circuit Court of Appeals reversed. Court membership"

Re: Judge said Meta illegally used books to build its AI

#292

Earlier quoted context omitted.

The lines for humans aren't clearly drawn, but they are drawn. The main difference is that humans are humans and LLMs are computer programs. I see no reason why we should even entertain the idea of extending human rights to computer programs, and so far, nobody has been able to give me any good reasons why. Furthermore, why are we only entertaining the human rights that can be used for profit-driven purposes? Why do…

You're thinking about it using the wrong framework IMO. It's not about the program's rights, it's about the human's rights to use the program. Not the machine's right to do something, but the human's right to do something through a machine, or make a machine do something.

No, because the entire argument hinges on the fact that LLMs learn, which is like humans learning, so it's transformative. That only works if you consider learning or transformation to be something that does not rely on the human spirit. Which, actually, most people do not believe. And it's pretty difficult to argue - we don't even know how learning works for people.

A lot of people just jump to LLMs learning like it's a foregone conclusion. Mm... no. You need to convince people of that. You'll find if you talk to non-tech people, they're not just going to believe you when you say that.

Why isn't an LLM more akin to a database or a compression algorithm? Why is it closer to human learning? After all, humans are humans and we have the exclusive right and power to determine what is human and what isn't. And database and compression algorithms are computer programs, of the same kind as an LLM.

Re: Judge said Meta illegally used books to build its AI

#293

Earlier quoted context omitted.

“Copy” is ambiguous here. Of course data is copied during training. That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data.

> That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data. While a tool being used to create infringing copies of some other work (whether or not it is the source material used to create the tool, and whether or not the infringing material is also verbatim copies) is relevant to whether the tool vendor is liable for contributory infringement for the infringing use of t…

> LLMs specifically, have been shown to have the capacity to make such copies

Exactly. I asked my Gemma how long of a quote it could give me of a given book, if I were the author & gave express permission, and I was a bit surprised it readily admitted it could

> Without Permission (Current Limit): Single sentence.

> With Broad Permission (Full Reproduction Allowed): I could theoretically quote the entire book.

Eye-opening (for me, at least).

Re: Judge said Meta illegally used books to build its AI

#294

Earlier quoted context omitted.

The output of the LLM is very different from the original, though. It’s hard to look at this and claim it isn’t.

In the general case, yes, but they can verifiably reproduce at least some copyrighted works verbatim, which implies, at the minimum, that their content is stored in model weights in some fashion.

Everyone knows the training data is stored in some way in the LLM. The point is the use of the copyrighted material is transformative. Remember google books, it literally shows photocopy of pages of books but the court ruled it’s fair use. A simplified explanation is book vs search engine and book vs ai chatbot are very different from each other.

Re: Judge said Meta illegally used books to build its AI

#295
I simply don't think that the copyright IP framework as it exists can be applied to training on this scale. Or, if it can, the relative value of any specific author/content creator's work is deminimis.

When the scale is a significant portion of all human text output ever, I don't think we're in the realm of any prior model. This is now something closer to how society attempts to approach natural resources like land, frequency bands, utility right-of-way, etc. I think this is the direction that laws and legislation should look to go. Or maybe not, I don't claim to have the answer, only that existing models are inadequate.

Re: Judge said Meta illegally used books to build its AI

#296
post #283

Earlier quoted context omitted.

Even if your belief that only the person *providing* the content is liable, do you honestly think a single person found all the content, downloaded it, directly trained the model themselves, and then deleted the content? If at any step the content was given or shared to anyone else for any reason, have they not converted into a provider themselves?

That in itself is a complicated issue, but I would suspect that this does not count as copyright infringement. The entity in possession of the data does not change, the copies and manipulation are performed by employees but at no time do they own what they are manipulating. If it were true that this constitutes a transfer of possession, then the targets for the lawsuit should be the individual employees, and I don't…

I would counter that they were operating under the direction of the corporation and, as such, the corporation should be liable. But, I'm not a lawyer.

My overall point is just that a corporation willfully violating copyright should be treated wholly differently than an individual.

I still think that, regardless of any copyright violation by the corporation in acquiring the content, that training with the data doesn't (or shouldn't) somehow taint the resulting model.

Re: Judge said Meta illegally used books to build its AI

#297
post #270

Earlier quoted context omitted.

> that's an exact copy. Not for the purposes of copyright law. > is that humans don't have separate RAM [or disk] And that turns out to be incredibly important. Humans can't create a lasting, shareable copy of a copyrighted work by consuming it.

Sure they can. You can learn a copyrighted work by hard, even indirectly, then quickly duplicate it by hand. Mozart was originally famous for making a business out of that.

> then quickly duplicate it by hand

And that's a copyright violation.

Re: Judge said Meta illegally used books to build its AI

#298

Earlier quoted context omitted.

Do you have any case law (other than Tivo or VHS time-shifting) that relates directly to books?

There was a relatively famous Google case regarding their digitization of books without the authors consent in 2015. Although it's not a perfectly analogous to this situation. In Googles case they were digitizing the books (that they did not own), and publishing snippets for search users to help them find books and other material that weren't indexed on the web. The court found they had that right, but did place some…

According to your link and this comment https://news.ycombinator.com/item?id=43899406 Google's scanning project was ruled fair use because it was "transformative" and didn't harm the market for the works. It allowed searching within books that was otherwise impossible at the time, but by not providing the full text of the books it didn't meaningfully reduce sales.

Someone photocopying a book to read on the toilet (and leave the original on their nightstand) isn't engaging in transformative use. They're also harming the market for the work because if they hadn't made this photocopy, they would've had to buy a second copy of the book to get the same benefit.

Re: Judge said Meta illegally used books to build its AI

#299
post #283

Earlier quoted context omitted.

That in itself is a complicated issue, but I would suspect that this does not count as copyright infringement. The entity in possession of the data does not change, the copies and manipulation are performed by employees but at no time do they own what they are manipulating. If it were true that this constitutes a transfer of possession, then the targets for the lawsuit should be the individual employees, and I don't…

I would counter that they were operating under the direction of the corporation and, as such, the corporation should be liable. But, I'm not a lawyer. My overall point is just that a corporation willfully violating copyright should be treated wholly differently than an individual. I still think that, regardless of any copyright violation by the corporation in acquiring the content, that training with the data doesn't…

>I still think that, regardless of any copyright violation by the corporation in acquiring the content, that training with the data doesn't (or shouldn't) somehow taint the resulting model.

Yes. Building the data set, acquiring the data set, training the model, using the model, and using the output of a model are all separate actions, it is not necessarily the case that any legal standing in one act transfers down the chain.

It's certainly not going to be resolved by a single case.

Re: Judge said Meta illegally used books to build its AI

#300
post #255

Earlier quoted context omitted.

> If the US decides to unilaterally shut down LLMs, that just means that the rest of the world will route around us. You're talking as if they are some kind of nationalized or publically-owned asset, as opposed to a bunch of for-profit, privately-owned silos.

These cases can set precedents that basically shut down all all future "useful" AI implementations if the judges go too far. You can bet that CCP doesn't care one whit about US copyright law if leads to them leapfrogging us, it will definitely be j-johna-jameson-laugh.gif . I think that's the point of the commenter from 20000ft

> if leads to them leapfrogging us

They won't be leapfrogging 'us'.

They'll be leapfrogging some privately owned, for-profit business.

Frankly, I don't give two shits about that, and I fail to see why anyone who doesn't own one of them should give one.

If they want me to have skin in the game, they should share either the profits, or the models. Until they do, it's not 'us', it's 'them'.

Post reply on HN