Live data from Hacker News

Judge said Meta illegally used books to build its AI

wired.com

301–310 of 352 posts

Re: Judge said Meta illegally used books to build its AI

#301
post #5

The title for this submission is somewhat misleading. The judge didn't make any sort of ruling, this is just reporting on a pretrial hearing. He also doesn't seem convinced as to how relevant downloading books from LibGen is to the case: > At times, it sounded like the case was the authors’ to lose, with [Judge] Chhabria noting that Meta was “destined to fail” if the plaintiffs could prove that Meta’s tools created s…

The RIAA lawyers never had to demonstrate that copying a DVD cratered the sales of their clients. They just got high penalties for infringers almost by default. Now that big capital wants to steal from individuals, big capital wins again. (Unrelatedly, has Boies ever won a high profile lawsuit? I remember him from the Bush/Gore recount issue, where he represented the Democrats.)

> The RIAA lawyers never had to demonstrate that copying a DVD cratered the sales of their clients.

Did any of the defendants raise a fair use defense based on a transformative use that they were making of the downloaded copies? If not, you are in the domain of "unlike legal situations lead to unlike decisions" which is not exactly surprising.

Re: Judge said Meta illegally used books to build its AI

#302

Earlier quoted context omitted.

The RIAA lawyers never had to demonstrate that copying a DVD cratered the sales of their clients. They just got high penalties for infringers almost by default. Now that big capital wants to steal from individuals, big capital wins again. (Unrelatedly, has Boies ever won a high profile lawsuit? I remember him from the Bush/Gore recount issue, where he represented the Democrats.)

> The RIAA lawyers never had to demonstrate that copying a DVD cratered the sales of their clients. They just got high penalties for infringers almost by default. The argument for 'fair use' in DVD copying/sharing is much weaker since the thing being shared in that case is a verbatim, digital copy of the work. 'Format shifting' is a tenuous argument, and it's pretty easily limited to making (and not distributing) per…

. If you analogize to human leaning

You can't make valid legal analogies to human learning when dealing with copyright law, because human brains are not a fixed media under copyright law, thus impressions in human brains are not copies of any kind under copyright law, thus "well, when you make an impression in the human brain, its not a copyright violation" is never a good legal analogy for when you make something that is in a form which can be a copy under copyright law.

Re: Judge said Meta illegally used books to build its AI

#303

Earlier quoted context omitted.

You're thinking about it using the wrong framework IMO. It's not about the program's rights, it's about the human's rights to use the program. Not the machine's right to do something, but the human's right to do something through a machine, or make a machine do something.

No, because the entire argument hinges on the fact that LLMs learn, which is like humans learning, so it's transformative. That only works if you consider learning or transformation to be something that does not rely on the human spirit. Which, actually, most people do not believe. And it's pretty difficult to argue - we don't even know how learning works for people. A lot of people just jump to LLMs learning like it…

Then the results would be the same, and it would still be fair use. I have yet to see an example that demonstrates LLMs plagiarize by default or by tendency.

Your causality seems to be inverted here. You seem to be implying that "learning" (or the ingestion and retention of information for the same means) is banned by default for everything, but we decide to allow it for humans as the sole exception. This is not the case. Everything not prohibited is allowed, and "intermediate copies" are considered to be vital to fair use by the court system.

Re: Judge said Meta illegally used books to build its AI

#304
post #58
post #54

Earlier quoted context omitted.

I'm curious what you mean by "in it's modern form". You seem to suggest there was a previous form that was not invented by corporations, but I don't believe that is the case.

Copyright was first established by governments and the Church prior to the invention of corporations.

> Copyright was first established by governments and the Church prior to the invention of corporations.

No, it wasn't. The first copyright law is generally held to be the Statute of Anne (1710) in Britain (the Licensing of the Press Act of 1662 which preceded it was not a "copyright law" in the sense that it did not provide for ownership of a right to publish/print particular works, but provided, as the name suggested, for licensing of printing presses, and revocation of said licenses -- it was more of a general censorship law); and even if you limit the scope to Britain (well, in either case, England before the Union with Scotland) corporations go back further (depending on whether you mean corporations as "entities which are not natural persons with legal personality"—the City of London Corporation's establishment is literally lost to history, but known to predate the Norman Conquest—or "joint stock companies"—the Company of Merchant Adventurers to New Lands in 1551 would be the first.)

Re: Judge said Meta illegally used books to build its AI

#305
post #232

Title seems misleading after reading the article.

It's mind blowing to me that the court might deny the right of the authors to control licensing of this kind of usage of their work.

If I put something up for anyone to read on the internet. And someone reads it on the internet. I can't really control that, right?

Now, if someone makes an infringing use of the thing I put up on the internet. Then I have some kind of recourse, at least through the courts, if I have a lot of money to pay lawyers.

But if someone makes a fair use of the thing I put up on the internet, then I don't have any recourse, because that's the way the law works.

As far as I understand it, using data as input data to a machine learning model that substantially transforms and does not duplicate the input data is currently believed to be fair use.

So, the training use of freely available data seems pretty straightforward that authors can't control when they make it freely available.

It seems like Facebook made use of data that wasn't freely available, though -- ebook rip library type stuff. That's the bit I think they could be in trouble for. But that's just a plain-old "Napster" style copyright question, as far as I understand it.

The lawyer's argument that Llama "obliterates the market" for written works seems weak. I, and anyone I know, put down AI slop fiction before the first paragraph is done, because it's not the same thing as real fiction.

Re: Judge said Meta illegally used books to build its AI

#306

Earlier quoted context omitted.

If your "fair use" is - of a commercial nature; - plagiarism; - substantially large (e.g. whole work); you're not on good legal footing.

I mean...sure? But that's not what you were saying. You said, as a blanket statement, that fair use requires citing the passage, which is not true. Fair use and plagiarism are related , but they are two separate things. Especially when talking about the legalities of things, as we are here, it's vital to be clear and accurate about what specific legal issues are under discussion. Facebook isn't claiming fair use to w…

Sure; say someone comes up with a great proverb in their book, or a nice phrase, and it catches on. People don't have to cite where it came from in order to use it. What if someone copies a half page paragraph? They are in a gray area where it will likely be in their favor if they can avoid the accusation of plagiarism, if they intend to use a fair use defense.

Anyway, the organizations behind popular LLMS are perpetrating massive copyright infringement and plagiarism. A fair use defense is rationally not possible.

Firstly, to claim fair use while you are making commercial use of the material is going to be difficult right off the bat. Fair use claims are bolstered by non-commercial use.

Small size of excerpts favors fair use claims. Can't do that here.

Transformative use? Not really; LLMs spit out the information verbatim if prompted in the right ways. Exploitative uses will probably not be viewed as transformative by the court.

Re: Judge said Meta illegally used books to build its AI

#307

Earlier quoted context omitted.

> The RIAA lawyers never had to demonstrate that copying a DVD cratered the sales of their clients. They just got high penalties for infringers almost by default. The argument for 'fair use' in DVD copying/sharing is much weaker since the thing being shared in that case is a verbatim, digital copy of the work. 'Format shifting' is a tenuous argument, and it's pretty easily limited to making (and not distributing) per…

> there's clearly no copyright infringement in a human learning from someone's work and creating their own output, even if it "copies" an artist's style or draws inspiration from someone's plot-line. What do you mean here by "clearly?" This is not at all clear, and court cases have been decided in the opposite direction. This case: https://www.reuters.com/article/lifestyle/marvin-gaye-family... is as far from what yo…

In the case you linked, it wasn't infringement for Robin Thicke to be impressed by Marvin Gaye. It was infringement to create a substantial copy.

Applying this same legal doctrine to LLMs, it's totally fine for them to train on copyrighted works. A problem would only arise if they produced substantial copies of those works.

It may be difficult to quantify the harm done in that case. Robin Thicke sold millions of copies of blurred lines. Each LLM output is generally consumed by a single person. There could be millions of examples of blatant copyright infringement by LLMs that will never come to light in court. There also could be vanishingly few. We may need new regulations to police it.

Re: Judge said Meta illegally used books to build its AI

#309

Earlier quoted context omitted.

No, because the entire argument hinges on the fact that LLMs learn, which is like humans learning, so it's transformative. That only works if you consider learning or transformation to be something that does not rely on the human spirit. Which, actually, most people do not believe. And it's pretty difficult to argue - we don't even know how learning works for people. A lot of people just jump to LLMs learning like it…

Then the results would be the same, and it would still be fair use. I have yet to see an example that demonstrates LLMs plagiarize by default or by tendency. Your causality seems to be inverted here. You seem to be implying that "learning" (or the ingestion and retention of information for the same means) is banned by default for everything, but we decide to allow it for humans as the sole exception. This is not the…

No, it wouldn't. Because if I record "Revenge of the Sith", compress it, and then distribute it for free online, that's obviously not fair use.

Fair use is pretty complicated. Part of Fair Use is the "The Effect of the Use on the Potential Market for or Value of the Work", which already puts even human commercial endeavors in a tough spot. You can make it work, but you have to really try. Satire like Weird Al or whatever isn't competing with the music it's satirizing, the venn diagram between those markets barely overlap. But a lot of LLM use cases are explicitly meant to obsolesce and siphon value from the things they used.

Like, why go to Getty Images when you could instead go to the glorified database, which has ingested all of Getty Images, and acquire an indistinguishable stock photo for free?

The only reason we're even really entertaining this is because people continually draw parallels to humans. You see, it's not stealing from Getty. It's more like if someone saw Getty Images and then went out and took a photo in that same flat, boring style. Except nobody saw anything. And nobody went out an took a photo.

Re: Judge said Meta illegally used books to build its AI

#310
post #223

Earlier quoted context omitted.

It's the same as a book or a film review, you can't get the film or the book back from it but the original material is still needed to produce it. Needing the original material isn't enough for claiming copyright infringement as we have existing counter examples

Movie reviews are fair use because they don't compete with the original work.

The AI models also do not compete with the original work.

Nobody is going to try to extract pages by pages a book from ChapGPT, let's be realistic. (and you can't anyways)

Post reply on HN