Live data from Hacker News

Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

understandingai.org

51–60 of 326 posts

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#51
post #10

It will generate a correct next token 42% of the time when prompted with a 50 token quote. Not 42% of the book. It's a pretty big distinction.

This means that if we start with 50% of the book then there is 42% chance that we can recreate the remaining 50%. What is the distinction between understanding and memorization? What is the chance that understanding results in memorization (may be in case of humans)?

[deleted]

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#52
post #46

Earlier quoted context omitted.

Sure, why not? lol https://www.reddit.com/r/DataHoarder/comments/1entowq/i_made... https://github.com/shloop/google-book-scraper The fact that Meta torrented Books3 and other datasets seems to be by self-admission by Meta employees who performed the work and/or oversaw those who themselves did the work, so that is not really under dispute or ambiguous. https://torrentfreak.com/meta-admits-use-of-pirated-book-dat...

Books3 was used in Llama1. We don't know if they used it later on.

My comparison was illustrative and analogous in nature. The copyright cartel is making a fruit of the poisonous tree type of argument. Whatever Meta are doing with LLMs is doing the heavy lifting that parity files used to do back in the Usenet days. I wouldn’t be surprised if BitTorrent or other similar caching and distribution mechanisms incorporate AI/LLMs to recognize an owl on the wire, draw the rest just in time in transit, and just send the diffs, or something like that.

The pictures are the same. All roads lead to Rome, so they say.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#53

I really wish we could get rid of copyright. It's going to hold us back long term.

I do too. But in the meantime, as long as it continues being used against anyone, it should be applied fairly. As long as anyone has to respect software licenses, for instance, then AIs should too. It doesn't stop being a problem just because it's done at larger scale.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#54

It will generate a correct next token 42% of the time when prompted with a 50 token quote. Not 42% of the book. It's a pretty big distinction.

next _50_ tokens 42% of the time

not just next token.

This is like: tell it a random sentence in the book, it will give you the next sentence 42% of time.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#55
I think it's important to recognize here that fanfiction.net has 850 thousand distinct pieces of Harry Potter fanction on it. Fifty thousand of which are more than 40k words in length. Many of which (no easy way to measure) directly reproducing parts of the original books.

archiveofourown.org has 500 thousand, some, but probably not the majority, of that are duplicated from fanfiction.net. 37 thousand of these are over 40 thousand words.

I.e. harry potter and its derivatives presumably appear a million times in the training set, and its hard to imagine a model that could discuss this cultural phenomena well without knowing quite a bit about the source material.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#56

I really wish we could get rid of copyright. It's going to hold us back long term.

We cannot get ride of it without finding a way to pay the creators that generate copyrighted works. I’m personally more in favor of significantly reducing the length of the copy right. I think 20-30 years is an interesting range. Artist get roughly a career length of time to profit off their creations, but there is much less incentive for major corporations to buy and horde IP.

We barely pay creators as it is for generating copyrighted works. Nearly every copywritten work is available on the internet, for free, right now. And creators are still getting paid, albeit poorly, but that's a constant throughout history.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#57
post #22

It's important to note the way it was measured: > the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time. So 50 tokens is not really…

Fair use is not a thing in every jurisdiction. In Germany for example there are cases where three words („wir sind Papst“) fall under copyright.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#59
post #8

From a quick web search I can find that there are book review sites that allow users to enter and rate verbatim "quotes" from books. This one [1] contains ~2000 [2] portions of a sentence, a paragraph or several paragraphs of Harry Potter and the Sorcerer's Stone. Could it be plausible that an LLM had ingested parts of the book via scrapping web pages like this and not the full copyrighted book and get results simila…

Meta has trained on LibGen so we don't really need to speculate.

https://www.wired.com/story/new-documents-unredacted-meta-co...

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#60
post #25

Earlier quoted context omitted.

> the physical encoding which definitely exists in my brain is a copyright violation First of all, we don't really know how the brain works. I get that you're being a snarky physicalist, but there's plenty of substance dualists, panpsychsts, etc. out there. So, some might say, this is a reductive description of what happens in our brains. Second of all, yes, if you tried to publish Harry Potter ( even if it was from…

Right but the physical encoding already exists in my brain or how can I reproduce it in the first place? We may not know how the encoding works but we do know that an encoding exists because a decoding is possible. My question is… is that in itself a violation of copyright? If not then as long as LLMs don’t make a publication it shouldn’t be a copyright violation right? Because we don’t understand how it’s encoded in…

To me the primary difference between the potential "copy" that exists in your brain and a potential "copy" that exists in the LLM, is that you can't make copies and distribute your brain to billions of people.

If you compressed a copy of HP as a .rar, you couldn't read that as is, but you could press a button and get HP out of it. To distribute that .rar would clearly be a copyright violation.

Likewise, you can't read whatever of HP exists in the LLM model directly, but you seemingly can press a bunch of buttons and get parts of it out. For some models, maybe you can get the entire thing. And I'm guessing you could train a model whose purpose is to output HP verbatim and get the book out of it as easily as de-compressing a .rar.

So, the question in my mind is, how similar is distributing the LLM model, or giving access to it, to distributing a .rar of HP. There's likely a spectrum of answers depending on the LLM

Post reply on HN