Live data from Hacker News

Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

understandingai.org

201–210 of 326 posts

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#201
post #109

Earlier quoted context omitted.

[flagged]

Repeating half of the book verbatim is not nearly the same as repeating a line.

If you prompt the LLM to output a book verbatim, then you violated the copyright, not the LLM. Just like if you take a book to a copier and make a copy of it, you are violating the copyright, not Xerox.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#202
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

> let's not pretend that an LLM that autocompletes a couple lines from harry potter with 50% accuracy is some massive new avenue to piracy. No one is using this as a substitute for buying the book.

You are completely missing the point. Have you read the actual article, because piracy isn't mention a single time.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#203

Well, so can a nontrivial number of people. It's Harry Potter we're talking about - it's up there with The Bible in popularity ranking. I'm gonna bet that Llama 3.1 can recall a significant portion of Pride and Prejudice too. With examples of this magnitude, it's normal and entirely expected this can happen - as it does with people[0] - the only thing this is really telling us is that the model doesn't understand its…

I keep waiting for the day when software stops being compared to a human person (a being with agency, free will, consciousness, and human rights of its own) for the purposes of justifying IP law circumvention.

Yes, there is no problem when a person reads some book and recalls pieces[0] of it in a suitable context. How would that in any way address when certain people create and distribute commercial software, providing it that piece as input, to perform such recall on demand and at scale, laundering and/or devaluing copyright, is unclear.

Notably, the above is being done not just to a few high-profile authors, but to all of us no matter what we do (be it music, software, writing, visual art).

What’s even worse, is that imaginably they train (or would train) the models to specifically not output those things verbatim specifically to thwart attempts to detect the presence of said works in training dataset (which would naturally reveal the model and its output being a derivative work).

Perhaps one could find some way of justifying that (people justified all sorts of stuff throughout history), but let it be something better than “the model is assumed to be a thinking human when it comes to IP abuse but unthinking tool when it comes to using it for personal benefit”.

[0] Of course, if you find me a single person on this planet capable of recalling 42% of any Harry Potter book, I’d be very impressed if I ever believed it.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#204

That's a clickbait title. What they are actually saying: Given one correct quoted sentence, the model has 42% chance of predicting the next sentence correctly. So, assuming you start with the first sentence and tell it to keep going, it has a 0.42^n odds of staying on track, where n is the n-th sentence. It seems to me, that if they didn't keep correcting it over and over again with real quotes, it wouldn't even get…

What would be a better title? You're correct that the title isn't accurate, however, click bait? I wouldn't say so. But I'm lacking imagination to find a better one. Interested to hear your suggestion.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#205

That's a clickbait title. What they are actually saying: Given one correct quoted sentence, the model has 42% chance of predicting the next sentence correctly. So, assuming you start with the first sentence and tell it to keep going, it has a 0.42^n odds of staying on track, where n is the n-th sentence. It seems to me, that if they didn't keep correcting it over and over again with real quotes, it wouldn't even get…

What would be a better title? You're correct that the title isn't accurate, however, click bait? I wouldn't say so. But I'm lacking imagination to find a better one. Interested to hear your suggestion.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#206
post #199
post #184

Earlier quoted context omitted.

It seems like a pretty reasonable argument and easy enough to make. A human with a great memory could probably recreate some absurd % of Harry Potter after reading it, there are some very unusual minds out there. It is clear that if they read Harry Potter and being capable of reproducing it on demand as a party trick that would be fair use. So the LLM should also be fair use since it is using a mechanism similar enou…

> is clear that if they read Harry Potter and reproduce it on demand as a party trick that would be fair use. Actually no that could be copyright infringement. Badly signing a recent pop song in public also qualifies as copyright infringement. Public performances count as copying here.

> Badly signing a recent pop song in public also qualifies as copyright infringement

For commercial purposes only. If someone sells a recreation of the Harry Potter book, it’s illegal regardless whether it was by memory, directly copying the book, or using an LLM. It’s the act of broadcasting it that’s infringing on copyright, not the content itself.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#207
post #199
post #184

Earlier quoted context omitted.

It seems like a pretty reasonable argument and easy enough to make. A human with a great memory could probably recreate some absurd % of Harry Potter after reading it, there are some very unusual minds out there. It is clear that if they read Harry Potter and being capable of reproducing it on demand as a party trick that would be fair use. So the LLM should also be fair use since it is using a mechanism similar enou…

> is clear that if they read Harry Potter and reproduce it on demand as a party trick that would be fair use. Actually no that could be copyright infringement. Badly signing a recent pop song in public also qualifies as copyright infringement. Public performances count as copying here.

Ah sorry. I mistyped. Being able to do that it would be fair use. I went back and fixed the comment.

Although frankly, as has been pointed out many times, the law is also stupid in what it prohibits and that should be fixed first as a priority. Its done some terrible damage to our culture. My family used to be part of a community choir until it shut down basically for copyright reasons.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#208
post #106

Earlier quoted context omitted.

Indeed but since when is a blatantly derived work only using 50% of a copyrighted work without permission a paragon of copyright compliance? Music artists get in trouble for using more than a sample without permission — imagine if they just used 45% of a whole song instead… I’m amazed AI companies haven’t been sued to oblivion yet. This utter stupidity only continues because we named a collection of matrices “Artific…

Music artists get in trouble for using more than a sample from other music artists without permission because their work is in direct competition with the work they're borrowing from. A ZIP file of a book is also in direct competition of the book, because you could open the ZIP file and read it instead of the book. A model that can take 50 tokens and give you a greater than 50% probability for the 50 next tokens 42%…

LLMs aren't probabilistic. The randomness is bolted on top by the cloud providers as a trick to give them a more humanistic feel.

Under the hood they are 100% deterministic, modulo quantization and rounding errors.

So yes, it is very much possible to use LLMs as a lossy compressed archive for texts.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#209

Do LLMs have any perception that Harry Potter is fiction or is it possible that they will give some magical advice based on fiction works that they have been trained with? edit: never mind, I’ll just ask ChatGPT

LLMs don't have "perception" at all, they only ever output a likely text completion token.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#210

Earlier quoted context omitted.

So if I memorized Harry Potter the physical encoding which definitely exists in my brain is a copyright violation?

I think humans get a special exception in cases like this

No they don't. Commercial intent is what is prosecuted in IP law.
Post reply on HN