It's important to note the way it was measured: > the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time. So 50 tokens is not really…
Fair use is not a thing in every jurisdiction. In Germany for example there are cases where three words („wir sind Papst“) fall under copyright.
Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
111–120 of 326 posts
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#112If the book was obtained legitimately, letting an LLM read it is not an issue.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#113Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#114Earlier quoted context omitted.
I think the argument is less about piracy and more that the model(s output) is a derivative work of Harry Potter, and the rights holder should be paid accordingly when it’s reproduced.
If the assertion in the parent comment is correct "nobody is using this as a substitute to buying the book" why should the rights holders get paid?
Repeat for every copyrighted work and you end up with publishers reasonably arguing meta would not be able to produce their LLM without copyrighted work, which they did not pay for.
It's an argument for the courts, of course.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#115On the other hand, it’s surprising that Llama memorized so much of Harry Potter and the Sorcerer's Stone. It's sold 120 million copies over 30 years. I've gotta think literally every passage is quoted online somewhere else a bunch of times. You could probably stitch together the full book quote-by-quote.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#116It's important to note the way it was measured: > the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time. So 50 tokens is not really…
All this study really says, is that models are really good at compressing the text of Harry Potter. You can't get Harry Potter out of it without prompting it with the missing bits - sure, impressively few bits, but is that surprising, considering how many references and fair use excerpts (like discussion of the story in public forums) it's seen? There's also the question of how many bits of originality there actually…
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#117I really wish we could get rid of copyright. It's going to hold us back long term.
We cannot get ride of it without finding a way to pay the creators that generate copyrighted works. I’m personally more in favor of significantly reducing the length of the copy right. I think 20-30 years is an interesting range. Artist get roughly a career length of time to profit off their creations, but there is much less incentive for major corporations to buy and horde IP.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#118I really wish we could get rid of copyright. It's going to hold us back long term.
We cannot get ride of it without finding a way to pay the creators that generate copyrighted works. I’m personally more in favor of significantly reducing the length of the copy right. I think 20-30 years is an interesting range. Artist get roughly a career length of time to profit off their creations, but there is much less incentive for major corporations to buy and horde IP.
As a full-time professional musician, I'm convinced I'll benefit much more from its deprecation than continuing to flog it into posterity. I don't think I know any musicians who believe that IP is career-relevant for them at this point.
(Granted, I play bluegrass, which has never fit into the copyright model of music in the first place)
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#119Earlier quoted context omitted.
I think the argument is less about piracy and more that the model(s output) is a derivative work of Harry Potter, and the rights holder should be paid accordingly when it’s reproduced.
If the assertion in the parent comment is correct "nobody is using this as a substitute to buying the book" why should the rights holders get paid?
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#120Earlier quoted context omitted.
All this study really says, is that models are really good at compressing the text of Harry Potter. You can't get Harry Potter out of it without prompting it with the missing bits - sure, impressively few bits, but is that surprising, considering how many references and fair use excerpts (like discussion of the story in public forums) it's seen? There's also the question of how many bits of originality there actually…
The alternate here is that Harry Potter is written with sentences that match the typical patterns of English and so, when you prompt with a part of the text, the LLM can complete it with above-random accuracy