Live data from Hacker News

Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

understandingai.org

21–30 of 326 posts

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#21

As I've said several times, the corpus is key: LLMs thus far "read" most anything, but should instead have well-curated corpora. "Garbage In, Garbage Out!(GIGO)" is the saying. While the Harry Potter series may be fun reading, it doesn't provide information about anything that isn't better covered elsewhere. Leave Harry Potter for a different "Harry Potter LLM". Train scientific LLMs to the level of a good early 20th…

> While the Harry Potter series may be fun reading, it doesn't provide information about anything that isn't better covered elsewhere It has copyright implications - if Claude can recollect 42% of a copyrighted product without attribution or royalties, how did Anthropic train it? > Train scientific LLMs to the level of a good early 20th century English major and then use science texts and research papers for the rema…

> It has copyright implications - if Claude can recollect 42% of a copyrighted product without attribution or royalties, how did Anthropic train it?

Personally I’m assuming the worst.

That being said, Harry Potter was such a big cultural phenomenon that I wonder to what degree might one actually be able to reconstruct the books based solely on publicly accessible derivative material.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#22
It's important to note the way it was measured:

> the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time

As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time.

So 50 tokens is not really very much, it's basically a sentence or two. Such a small amount would probably generally fall under fair use on its own. To allege a true copyright violation you'd still need to show that you can chain those together or use some other method to build actual substantial portions of the book. And if it only gets it right 50% of the time, that seems like it would be very hard to do with high fidelity.

Having said all that, what is really interesting is how different the latest Llama 70b is from previous versions. It does suggest that Meta maybe got a bit desperate and started over-training on certain materials that greatly increased its direct recall behaviour.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#23

Earlier quoted context omitted.

> While the Harry Potter series may be fun reading, it doesn't provide information about anything that isn't better covered elsewhere It has copyright implications - if Claude can recollect 42% of a copyrighted product without attribution or royalties, how did Anthropic train it? > Train scientific LLMs to the level of a good early 20th century English major and then use science texts and research papers for the rema…

So if I memorized Harry Potter the physical encoding which definitely exists in my brain is a copyright violation?

If you perform it from memory in public without paying royalties then yes, yes it is.

Should it be? Different question.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#25

Earlier quoted context omitted.

> While the Harry Potter series may be fun reading, it doesn't provide information about anything that isn't better covered elsewhere It has copyright implications - if Claude can recollect 42% of a copyrighted product without attribution or royalties, how did Anthropic train it? > Train scientific LLMs to the level of a good early 20th century English major and then use science texts and research papers for the rema…

So if I memorized Harry Potter the physical encoding which definitely exists in my brain is a copyright violation?

> the physical encoding which definitely exists in my brain is a copyright violation

First of all, we don't really know how the brain works. I get that you're being a snarky physicalist, but there's plenty of substance dualists, panpsychsts, etc. out there. So, some might say, this is a reductive description of what happens in our brains.

Second of all, yes, if you tried to publish Harry Potter (even if it was from memory), you would get in trouble for copyright violation.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#26

Earlier quoted context omitted.

> While the Harry Potter series may be fun reading, it doesn't provide information about anything that isn't better covered elsewhere It has copyright implications - if Claude can recollect 42% of a copyrighted product without attribution or royalties, how did Anthropic train it? > Train scientific LLMs to the level of a good early 20th century English major and then use science texts and research papers for the rema…

So if I memorized Harry Potter the physical encoding which definitely exists in my brain is a copyright violation?

Only if you charge someone to reproduce it for them

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#27
post #17

As I've said several times, the corpus is key: LLMs thus far "read" most anything, but should instead have well-curated corpora. "Garbage In, Garbage Out!(GIGO)" is the saying. While the Harry Potter series may be fun reading, it doesn't provide information about anything that isn't better covered elsewhere. Leave Harry Potter for a different "Harry Potter LLM". Train scientific LLMs to the level of a good early 20th…

That's got nothing to do with it. It's all about copyright. Can it reproduce its training data verbatim? If so, Meta is in hot water.

I read harry potter, and you ask me about a page, and I can recite it verbatim, did I just commit copyright infringement?

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#28
post #22

It's important to note the way it was measured: > the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time. So 50 tokens is not really…

> So 50 tokens is not really very much, it's basically a sentence or two. Such a small amount would probably generally fall under fair use on its own.

That’s what I was thinking as I read the methodology.

If they dropped the same prompt fragment into Google (or any search engine) how often would they get the next 50 tokens worth of text returned in the search results summaries?

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#29
post #22

It's important to note the way it was measured: > the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time. So 50 tokens is not really…

(Disclaimer: haven't read the original paper)

It sounds like a ridiculous way to measure it. Producing 50-token excerpts absolutely doesn't translate to "recall X percent of Harry Potter" for me.

(Edit: I read this article. Nothing burger if its interpretation of the original paper is correct.)

Post reply on HN