It will generate a correct next token 42% of the time when prompted with a 50 token quote. Not 42% of the book. It's a pretty big distinction.
This means that if we start with 50% of the book then there is 42% chance that we can recreate the remaining 50%. What is the distinction between understanding and memorization? What is the chance that understanding results in memorization (may be in case of humans)?
Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
51–60 of 326 posts
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#52Earlier quoted context omitted.
Sure, why not? lol https://www.reddit.com/r/DataHoarder/comments/1entowq/i_made... https://github.com/shloop/google-book-scraper The fact that Meta torrented Books3 and other datasets seems to be by self-admission by Meta employees who performed the work and/or oversaw those who themselves did the work, so that is not really under dispute or ambiguous. https://torrentfreak.com/meta-admits-use-of-pirated-book-dat...
Books3 was used in Llama1. We don't know if they used it later on.
The pictures are the same. All roads lead to Rome, so they say.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#53I really wish we could get rid of copyright. It's going to hold us back long term.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#54It will generate a correct next token 42% of the time when prompted with a 50 token quote. Not 42% of the book. It's a pretty big distinction.
not just next token.
This is like: tell it a random sentence in the book, it will give you the next sentence 42% of time.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#55archiveofourown.org has 500 thousand, some, but probably not the majority, of that are duplicated from fanfiction.net. 37 thousand of these are over 40 thousand words.
I.e. harry potter and its derivatives presumably appear a million times in the training set, and its hard to imagine a model that could discuss this cultural phenomena well without knowing quite a bit about the source material.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#56I really wish we could get rid of copyright. It's going to hold us back long term.
We cannot get ride of it without finding a way to pay the creators that generate copyrighted works. I’m personally more in favor of significantly reducing the length of the copy right. I think 20-30 years is an interesting range. Artist get roughly a career length of time to profit off their creations, but there is much less incentive for major corporations to buy and horde IP.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#57It's important to note the way it was measured: > the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time. So 50 tokens is not really…
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#58It will generate a correct next token 42% of the time when prompted with a 50 token quote. Not 42% of the book. It's a pretty big distinction.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#59From a quick web search I can find that there are book review sites that allow users to enter and rate verbatim "quotes" from books. This one [1] contains ~2000 [2] portions of a sentence, a paragraph or several paragraphs of Harry Potter and the Sorcerer's Stone. Could it be plausible that an LLM had ingested parts of the book via scrapping web pages like this and not the full copyrighted book and get results simila…
https://www.wired.com/story/new-documents-unredacted-meta-co...
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#60Earlier quoted context omitted.
> the physical encoding which definitely exists in my brain is a copyright violation First of all, we don't really know how the brain works. I get that you're being a snarky physicalist, but there's plenty of substance dualists, panpsychsts, etc. out there. So, some might say, this is a reductive description of what happens in our brains. Second of all, yes, if you tried to publish Harry Potter ( even if it was from…
Right but the physical encoding already exists in my brain or how can I reproduce it in the first place? We may not know how the encoding works but we do know that an encoding exists because a decoding is possible. My question is… is that in itself a violation of copyright? If not then as long as LLMs don’t make a publication it shouldn’t be a copyright violation right? Because we don’t understand how it’s encoded in…
If you compressed a copy of HP as a .rar, you couldn't read that as is, but you could press a button and get HP out of it. To distribute that .rar would clearly be a copyright violation.
Likewise, you can't read whatever of HP exists in the LLM model directly, but you seemingly can press a bunch of buttons and get parts of it out. For some models, maybe you can get the entire thing. And I'm guessing you could train a model whose purpose is to output HP verbatim and get the book out of it as easily as de-compressing a .rar.
So, the question in my mind is, how similar is distributing the LLM model, or giving access to it, to distributing a .rar of HP. There's likely a spectrum of answers depending on the LLM