Live data from Hacker News

Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

understandingai.org

101–110 of 326 posts

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#101
post #60

Earlier quoted context omitted.

To me the primary difference between the potential "copy" that exists in your brain and a potential "copy" that exists in the LLM, is that you can't make copies and distribute your brain to billions of people. If you compressed a copy of HP as a .rar, you couldn't read that as is, but you could press a button and get HP out of it. To distribute that .rar would clearly be a copyright violation. Likewise, you can't rea…

> that exists in the LLM, is that you can't make copies and distribute your brain to billions of people. I can record myself reciting the full Harry Potter book then distribute it on YouTube. Could do the exact same thing with an LLM. The potential for distribution exists in both cases. Why is one illegal and the other not?

> I can record myself reciting the full Harry Potter book then distribute it on YouTube

Not legally you can't. Both of your examples are copyright violations

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#102
post #46

Earlier quoted context omitted.

Books3 was used in Llama1. We don't know if they used it later on.

All of the major AI models these days use "clean" datasets stripped of copyrighted material. They also use data from the previous models, so I'm not sure how "clean" it really is

All written text is copyrighted, with few exceptions like court transcripts. I own the copyright to this inane comment. I sincerely doubt that all copyrighted material is scrubbed.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#103
post #75

Earlier quoted context omitted.

I think the argument is less about piracy and more that the model(s output) is a derivative work of Harry Potter, and the rights holder should be paid accordingly when it’s reproduced.

That may be relevant in the NYT vs OpenAI case, since NYT was supposedly able to reproduce entire articles in ChatGPT. Here Llama is predicting one sentence at a time when fed the previous one, with 50% accuracy, for 42% of the book. That can easily be written off as fair use.

> Here Llama is predicting one sentence at a time when fed the previous one, with 50% accuracy, for 42% of the book. That can easily be written off as fair use.

Is that fair use, or is that compression of the verbatim source?

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#104
post #99
post #90

Earlier quoted context omitted.

You really don't see the difference between Google indexing the content of third parties and directly hosting/distributing the content itself?

Hosting model weights is not hosting / distributing the content.

Of course it is.

It's just a form of compression.

If I train an autoencoder on an image, and distribute the weights, that would obviously be the same as distributing the content. Just because the content is commingled with lots of other content doesn't make it disappear.

Besides, where did the sections of text from the input works that show up in the output text come from? Divine inspiration? God whispering to the machine?

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#105
post #98
post #90

Earlier quoted context omitted.

You really don't see the difference between Google indexing the content of third parties and directly hosting/distributing the content itself?

Where are they putting any blame on Google here?

Where did I say they were?

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#106
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

Indeed but since when is a blatantly derived work only using 50% of a copyrighted work without permission a paragon of copyright compliance? Music artists get in trouble for using more than a sample without permission — imagine if they just used 45% of a whole song instead… I’m amazed AI companies haven’t been sued to oblivion yet. This utter stupidity only continues because we named a collection of matrices “Artific…

Music artists get in trouble for using more than a sample from other music artists without permission because their work is in direct competition with the work they're borrowing from.

A ZIP file of a book is also in direct competition of the book, because you could open the ZIP file and read it instead of the book.

A model that can take 50 tokens and give you a greater than 50% probability for the 50 next tokens 42% of the time is not in direct competition with the book, since starting from the beginning you'll lose the plot fairly quickly unless you already have the full book, and unlike music sampling from other music, the model output isn't good enough to read it instead of the book.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#107
post #89
post #74

Earlier quoted context omitted.

The article fails to mention or understand the volume of content here. Every, literally every, part of these books is quoted and "talked about" (in the sense of used in unlicensed derivative works). And yes, I read the article before commenting. I don't appreciate the baseless insinuation to the contrary.

Even assuming you are correct, which I'm skeptical of, does this make it better? It's essentially the same thing, they are copying from a source that is violating copyright, whether that's a pirated book directly or a pirated book via fanficton.

Generally I think it matters a great deal to get the facts right when discussing something with nuance.

Is this specific fact required to make my beliefs consistent... Yes I think it is, but if you disagree with me in other ways it might not be important to your beliefs.

Legally (note: not a lawyer) I'm generally of the opinion that

A) Torrenting these books was probably copyright infringement on Meta's part. They should have done so legally by scanning lawfully acquired copies like Google did with Google Books.

B) Everything else here that Meta did falls under the fair use and de minimis exceptions to copyrights prohibition on copying copyrighted works without a license.

And if it was copying significant amounts of a work that appeared only once in its training set into the model the de minimis argument would fall apart.

Morally I'm of the opinion that copyright law's prohibition on deeply interacting with our cultural artifacts by creating derivative works is incredibly unfair and bad for society. This extends to a belief that the communities that do this should not be excluded from technological developments because there entire existence is unjustly outlawed.

Incidentally I don't believe that browsing a site that complies with the DMCA and viewing what it lawfully serves you constitutes piracy, so I can't agree with your characterization of events either. The fanfiction was not pirated just because it was likely unlawful to produce in the US.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#108

Earlier quoted context omitted.

All of the major AI models these days use "clean" datasets stripped of copyrighted material. They also use data from the previous models, so I'm not sure how "clean" it really is

All written text is copyrighted, with few exceptions like court transcripts. I own the copyright to this inane comment. I sincerely doubt that all copyrighted material is scrubbed.

Your brief comment is hardly copyrightable. Which makes your point moot.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#109
post #104
post #99

Earlier quoted context omitted.

Hosting model weights is not hosting / distributing the content.

Of course it is. It's just a form of compression. If I train an autoencoder on an image, and distribute the weights, that would obviously be the same as distributing the content. Just because the content is commingled with lots of other content doesn't make it disappear. Besides, where did the sections of text from the input works that show up in the output text come from? Divine inspiration? God whispering to the ma…

[flagged]

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#110
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

I think the argument is less about piracy and more that the model(s output) is a derivative work of Harry Potter, and the rights holder should be paid accordingly when it’s reproduced.

If the assertion in the parent comment is correct "nobody is using this as a substitute to buying the book" why should the rights holders get paid?
Post reply on HN