Live data from Hacker News

Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

understandingai.org

61–70 of 326 posts

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#61

I can recall about 12% of the first Harry Potter book so it's interesting to see Llama is only 4x smarter than me. I will catch up.

How many r’s are there in strawberry?

There are 3 R's in strawberry just like in Harry Potter!

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#62

Earlier quoted context omitted.

I read harry potter, and you ask me about a page, and I can recite it verbatim, did I just commit copyright infringement?

Are you selling your ability to recite stuff? Then certainly.

there are plenty of open source LLMs trained on harry potter, is that fine?

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#63
post #22

It's important to note the way it was measured: > the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time. So 50 tokens is not really…

(Disclaimer: haven't read the original paper) It sounds like a ridiculous way to measure it. Producing 50-token excerpts absolutely doesn't translate to "recall X percent of Harry Potter" for me. (Edit: I read this article. Nothing burger if its interpretation of the original paper is correct.)

Their methodology seems reasonable to me.

To clarify, they look at the probability a model will produce a verbatim 50-token excerpt given the preceding 50 tokens. They evaluate this for all sequences in the book using a sliding window of 10 characters (NB: not tokens). Sequences from Harry Potter have substantially higher probabilities of being reproduced than sequences from less well-known books.

Whether this is "recall" is, of course, one of those tricky semantic arguments we have yet to settle when it comes to LLMs.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#64
As an experiment I searched Google for "harry potter and the sorcerer's stone text":

- the first result is a pdf of the full book

- the second result is a txt of the full book

- the third result is a pdf of the complete harry potter collection

- the fourth result is a txt of the full book (hosted on github funny enough)

Further down there are similar copies from the internet archive and dozens of other sites. All in the first 2-3 pages.

I get that copyright is a problem, but let's not pretend that an LLM that autocompletes a couple lines from harry potter with 50% accuracy is some massive new avenue to piracy. No one is using this as a substitute for buying the book.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#65
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

> let's not pretend that an LLM that autocompletes a couple lines from harry potter with 50% accuracy is some massive new avenue to piracy. No one is using this as a substitute for buying the book.

Well, luckily the article points out what people are actually alleging:

> There are actually three distinct theories of how training a model on copyrighted works could infringe copyright:

> Training on a copyrighted work is inherently infringing because the training process involves making a digital copy of the work.

> The training process copies information from the training data into the model, making the model a derivative work under copyright law.

> Infringement occurs when a model generates (portions of) a copyrighted work.

None of those claim that these models are a substitute to buying the books. That's not what the plaintiffs are alleging. Infringing on a copyright is not only a matter of privacy (piracy is one of many ways to infringe copyright)

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#66
post #60

Earlier quoted context omitted.

Right but the physical encoding already exists in my brain or how can I reproduce it in the first place? We may not know how the encoding works but we do know that an encoding exists because a decoding is possible. My question is… is that in itself a violation of copyright? If not then as long as LLMs don’t make a publication it shouldn’t be a copyright violation right? Because we don’t understand how it’s encoded in…

To me the primary difference between the potential "copy" that exists in your brain and a potential "copy" that exists in the LLM, is that you can't make copies and distribute your brain to billions of people. If you compressed a copy of HP as a .rar, you couldn't read that as is, but you could press a button and get HP out of it. To distribute that .rar would clearly be a copyright violation. Likewise, you can't rea…

> that exists in the LLM, is that you can't make copies and distribute your brain to billions of people.

I can record myself reciting the full Harry Potter book then distribute it on YouTube.

Could do the exact same thing with an LLM. The potential for distribution exists in both cases. Why is one illegal and the other not?

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#67
post #8

From a quick web search I can find that there are book review sites that allow users to enter and rate verbatim "quotes" from books. This one [1] contains ~2000 [2] portions of a sentence, a paragraph or several paragraphs of Harry Potter and the Sorcerer's Stone. Could it be plausible that an LLM had ingested parts of the book via scrapping web pages like this and not the full copyrighted book and get results simila…

This is in fact mentioned and addressed in the article. Also, there is pretty clear cut evidence Meta used pirated book data sets knowingly to train the earlier Llama models

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#68
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

I think the argument is less about piracy and more that the model(s output) is a derivative work of Harry Potter, and the rights holder should be paid accordingly when it’s reproduced.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#69
post #46

Earlier quoted context omitted.

Sure, why not? lol https://www.reddit.com/r/DataHoarder/comments/1entowq/i_made... https://github.com/shloop/google-book-scraper The fact that Meta torrented Books3 and other datasets seems to be by self-admission by Meta employees who performed the work and/or oversaw those who themselves did the work, so that is not really under dispute or ambiguous. https://torrentfreak.com/meta-admits-use-of-pirated-book-dat...

Books3 was used in Llama1. We don't know if they used it later on.

All of the major AI models these days use "clean" datasets stripped of copyrighted material.

They also use data from the previous models, so I'm not sure how "clean" it really is

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#70
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

Let's also not pretend that "massive new" is the only relevant issue
Post reply on HN