I can recall about 12% of the first Harry Potter book so it's interesting to see Llama is only 4x smarter than me. I will catch up.
How many r’s are there in strawberry?
Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
61–70 of 326 posts
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#62Earlier quoted context omitted.
I read harry potter, and you ask me about a page, and I can recite it verbatim, did I just commit copyright infringement?
Are you selling your ability to recite stuff? Then certainly.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#63It's important to note the way it was measured: > the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time. So 50 tokens is not really…
(Disclaimer: haven't read the original paper) It sounds like a ridiculous way to measure it. Producing 50-token excerpts absolutely doesn't translate to "recall X percent of Harry Potter" for me. (Edit: I read this article. Nothing burger if its interpretation of the original paper is correct.)
To clarify, they look at the probability a model will produce a verbatim 50-token excerpt given the preceding 50 tokens. They evaluate this for all sequences in the book using a sliding window of 10 characters (NB: not tokens). Sequences from Harry Potter have substantially higher probabilities of being reproduced than sequences from less well-known books.
Whether this is "recall" is, of course, one of those tricky semantic arguments we have yet to settle when it comes to LLMs.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#64- the first result is a pdf of the full book
- the second result is a txt of the full book
- the third result is a pdf of the complete harry potter collection
- the fourth result is a txt of the full book (hosted on github funny enough)
Further down there are similar copies from the internet archive and dozens of other sites. All in the first 2-3 pages.
I get that copyright is a problem, but let's not pretend that an LLM that autocompletes a couple lines from harry potter with 50% accuracy is some massive new avenue to piracy. No one is using this as a substitute for buying the book.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#65As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…
Well, luckily the article points out what people are actually alleging:
> There are actually three distinct theories of how training a model on copyrighted works could infringe copyright:
> Training on a copyrighted work is inherently infringing because the training process involves making a digital copy of the work.
> The training process copies information from the training data into the model, making the model a derivative work under copyright law.
> Infringement occurs when a model generates (portions of) a copyrighted work.
None of those claim that these models are a substitute to buying the books. That's not what the plaintiffs are alleging. Infringing on a copyright is not only a matter of privacy (piracy is one of many ways to infringe copyright)
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#66Earlier quoted context omitted.
Right but the physical encoding already exists in my brain or how can I reproduce it in the first place? We may not know how the encoding works but we do know that an encoding exists because a decoding is possible. My question is… is that in itself a violation of copyright? If not then as long as LLMs don’t make a publication it shouldn’t be a copyright violation right? Because we don’t understand how it’s encoded in…
To me the primary difference between the potential "copy" that exists in your brain and a potential "copy" that exists in the LLM, is that you can't make copies and distribute your brain to billions of people. If you compressed a copy of HP as a .rar, you couldn't read that as is, but you could press a button and get HP out of it. To distribute that .rar would clearly be a copyright violation. Likewise, you can't rea…
I can record myself reciting the full Harry Potter book then distribute it on YouTube.
Could do the exact same thing with an LLM. The potential for distribution exists in both cases. Why is one illegal and the other not?
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#67From a quick web search I can find that there are book review sites that allow users to enter and rate verbatim "quotes" from books. This one [1] contains ~2000 [2] portions of a sentence, a paragraph or several paragraphs of Harry Potter and the Sorcerer's Stone. Could it be plausible that an LLM had ingested parts of the book via scrapping web pages like this and not the full copyrighted book and get results simila…
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#68As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#69Earlier quoted context omitted.
Sure, why not? lol https://www.reddit.com/r/DataHoarder/comments/1entowq/i_made... https://github.com/shloop/google-book-scraper The fact that Meta torrented Books3 and other datasets seems to be by self-admission by Meta employees who performed the work and/or oversaw those who themselves did the work, so that is not really under dispute or ambiguous. https://torrentfreak.com/meta-admits-use-of-pirated-book-dat...
Books3 was used in Llama1. We don't know if they used it later on.
They also use data from the previous models, so I'm not sure how "clean" it really is
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#70As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…