Live data from Hacker News

Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

understandingai.org

111–120 of 326 posts

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#111
post #57
post #22

It's important to note the way it was measured: > the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time. So 50 tokens is not really…

Fair use is not a thing in every jurisdiction. In Germany for example there are cases where three words („wir sind Papst“) fall under copyright.

Germany does not have something called "fair use," but it does have provisions for uses that are fair. For example your use of the three words to talk about their copyrighted status is perfectly legal in Germany. That somebody wasn't allowed to use them in a specific way in the past doesn't mean that nobody is allowed to use them in any way.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#112
Many people could also produce text snippets from memory. I dispute that reading a book is a copyright violation. Copying and distributing a book, yes, but just reading it - no.

If the book was obtained legitimately, letting an LLM read it is not an issue.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#113
I'm surprised no one in the comments has mentioned overfitting. Perhaps this is too obvious but I think of it as a very clear bug in a model if it asserts something to be true because it has heard it once. I realize that training a model is not easy, but this is something that should've been caught before it was released. Either QA is sleeping on the job or they have intentionally released a model with serious flaws in its design/training. I also understand the intense pressure to release early and often, but this type of thing isn't a warning.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#114

Earlier quoted context omitted.

I think the argument is less about piracy and more that the model(s output) is a derivative work of Harry Potter, and the rights holder should be paid accordingly when it’s reproduced.

If the assertion in the parent comment is correct "nobody is using this as a substitute to buying the book" why should the rights holders get paid?

The argument is meta used the book so the LLM can be considered a derivative work in some sense.

Repeat for every copyrighted work and you end up with publishers reasonably arguing meta would not be able to produce their LLM without copyrighted work, which they did not pay for.

It's an argument for the courts, of course.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#115

On the other hand, it’s surprising that Llama memorized so much of Harry Potter and the Sorcerer's Stone. It's sold 120 million copies over 30 years. I've gotta think literally every passage is quoted online somewhere else a bunch of times. You could probably stitch together the full book quote-by-quote.

How many could do it from memory?

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#116
post #22

It's important to note the way it was measured: > the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time. So 50 tokens is not really…

All this study really says, is that models are really good at compressing the text of Harry Potter. You can't get Harry Potter out of it without prompting it with the missing bits - sure, impressively few bits, but is that surprising, considering how many references and fair use excerpts (like discussion of the story in public forums) it's seen? There's also the question of how many bits of originality there actually…

The alternate here is that Harry Potter is written with sentences that match the typical patterns of English and so, when you prompt with a part of the text, the LLM can complete it with above-random accuracy

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#117

I really wish we could get rid of copyright. It's going to hold us back long term.

We cannot get ride of it without finding a way to pay the creators that generate copyrighted works. I’m personally more in favor of significantly reducing the length of the copy right. I think 20-30 years is an interesting range. Artist get roughly a career length of time to profit off their creations, but there is much less incentive for major corporations to buy and horde IP.

[deleted]

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#118

I really wish we could get rid of copyright. It's going to hold us back long term.

We cannot get ride of it without finding a way to pay the creators that generate copyrighted works. I’m personally more in favor of significantly reducing the length of the copy right. I think 20-30 years is an interesting range. Artist get roughly a career length of time to profit off their creations, but there is much less incentive for major corporations to buy and horde IP.

I do not think it's creators that are the constituency holding up deprecation.

As a full-time professional musician, I'm convinced I'll benefit much more from its deprecation than continuing to flog it into posterity. I don't think I know any musicians who believe that IP is career-relevant for them at this point.

(Granted, I play bluegrass, which has never fit into the copyright model of music in the first place)

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#119

Earlier quoted context omitted.

I think the argument is less about piracy and more that the model(s output) is a derivative work of Harry Potter, and the rights holder should be paid accordingly when it’s reproduced.

If the assertion in the parent comment is correct "nobody is using this as a substitute to buying the book" why should the rights holders get paid?

The argument is whether the LLM training on the copyrighted work is Fair Use or not. Should META pay for the copyright on works it ingests for training purposes?

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#120

Earlier quoted context omitted.

All this study really says, is that models are really good at compressing the text of Harry Potter. You can't get Harry Potter out of it without prompting it with the missing bits - sure, impressively few bits, but is that surprising, considering how many references and fair use excerpts (like discussion of the story in public forums) it's seen? There's also the question of how many bits of originality there actually…

The alternate here is that Harry Potter is written with sentences that match the typical patterns of English and so, when you prompt with a part of the text, the LLM can complete it with above-random accuracy

Or else, LLMs show that copyright and IP are ridiculous concepts that should be abolished
Post reply on HN