Live data from Hacker News

Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

understandingai.org

31–40 of 326 posts

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#31

Earlier quoted context omitted.

> While the Harry Potter series may be fun reading, it doesn't provide information about anything that isn't better covered elsewhere It has copyright implications - if Claude can recollect 42% of a copyrighted product without attribution or royalties, how did Anthropic train it? > Train scientific LLMs to the level of a good early 20th century English major and then use science texts and research papers for the rema…

Why are you talking about Claude and Anthropic?

It’s not unreasonable to suspect they are doing the same. The article starts with a description of a lawsuit NY Times brought against OpenAI for similar reasons. The big difference is that research presented here is only possible with open weight models. OAI and Anthropic don’t make the base models available, so it’s easier to hide the fact that you’ve used copyrighted material by instruction post-training. And I’m not sure you can get the logprobs for specific tokens from their APIs either (which is what the researchers did to make the figures and come up with a concrete number like 42%)

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#32
On the other hand, it’s surprising that Llama memorized so much of Harry Potter and the Sorcerer's Stone.

It's sold 120 million copies over 30 years. I've gotta think literally every passage is quoted online somewhere else a bunch of times. You could probably stitch together the full book quote-by-quote.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#33
post #17

Earlier quoted context omitted.

That's got nothing to do with it. It's all about copyright. Can it reproduce its training data verbatim? If so, Meta is in hot water.

I read harry potter, and you ask me about a page, and I can recite it verbatim, did I just commit copyright infringement?

[deleted]

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#34

Earlier quoted context omitted.

> While the Harry Potter series may be fun reading, it doesn't provide information about anything that isn't better covered elsewhere It has copyright implications - if Claude can recollect 42% of a copyrighted product without attribution or royalties, how did Anthropic train it? > Train scientific LLMs to the level of a good early 20th century English major and then use science texts and research papers for the rema…

So if I memorized Harry Potter the physical encoding which definitely exists in my brain is a copyright violation?

The end of "Fahrenheit 451" set a horrible precedent. Damn you, Bradbury!

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#36
post #22

It's important to note the way it was measured: > the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time. So 50 tokens is not really…

This sounds almost like "Works every time (50% of the time)."

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#37

On the other hand, it’s surprising that Llama memorized so much of Harry Potter and the Sorcerer's Stone. It's sold 120 million copies over 30 years. I've gotta think literally every passage is quoted online somewhere else a bunch of times. You could probably stitch together the full book quote-by-quote.

But also we know for a fact that Meta trained their models on pirated books. So there's no need to invent a hare brained scheme of stitching together bits and pieces like that.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#38
post #25

Earlier quoted context omitted.

So if I memorized Harry Potter the physical encoding which definitely exists in my brain is a copyright violation?

> the physical encoding which definitely exists in my brain is a copyright violation First of all, we don't really know how the brain works. I get that you're being a snarky physicalist, but there's plenty of substance dualists, panpsychsts, etc. out there. So, some might say, this is a reductive description of what happens in our brains. Second of all, yes, if you tried to publish Harry Potter ( even if it was from…

Right but the physical encoding already exists in my brain or how can I reproduce it in the first place? We may not know how the encoding works but we do know that an encoding exists because a decoding is possible.

My question is… is that in itself a violation of copyright?

If not then as long as LLMs don’t make a publication it shouldn’t be a copyright violation right? Because we don’t understand how it’s encoded in LLMs either. It is literally the same concept.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#39

On the other hand, it’s surprising that Llama memorized so much of Harry Potter and the Sorcerer's Stone. It's sold 120 million copies over 30 years. I've gotta think literally every passage is quoted online somewhere else a bunch of times. You could probably stitch together the full book quote-by-quote.

Probably not?

Sure there are just ~75,000 words in HP1, and there are probably many times that amount in direct quotes online. However the quotes aren’t even distributed across the entire text. For every quote of charming the snake in a zoo there will be a thousand “you’re a wizard harry”, and those are two prominent plot points.

I suspect the least popular of all direct quotes from HP1 aren’t using the quotes in fair use, and are just replicating large sections of the novel.

Or maybe it really is just so popular that super nerds have quoted the entire novel arguing about the aspects of wand making, or the contents of every lecture.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#40
post #22

It's important to note the way it was measured: > the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time. So 50 tokens is not really…

Even if it is recalling it 50 tokens at a time, the half of the book is in some sense in there, right?
Post reply on HN