Earlier quoted context omitted.
[flagged]
Repeating half of the book verbatim is not nearly the same as repeating a line.
Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
201–210 of 326 posts
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#202As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…
You are completely missing the point. Have you read the actual article, because piracy isn't mention a single time.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#203Well, so can a nontrivial number of people. It's Harry Potter we're talking about - it's up there with The Bible in popularity ranking. I'm gonna bet that Llama 3.1 can recall a significant portion of Pride and Prejudice too. With examples of this magnitude, it's normal and entirely expected this can happen - as it does with people[0] - the only thing this is really telling us is that the model doesn't understand its…
Yes, there is no problem when a person reads some book and recalls pieces[0] of it in a suitable context. How would that in any way address when certain people create and distribute commercial software, providing it that piece as input, to perform such recall on demand and at scale, laundering and/or devaluing copyright, is unclear.
Notably, the above is being done not just to a few high-profile authors, but to all of us no matter what we do (be it music, software, writing, visual art).
What’s even worse, is that imaginably they train (or would train) the models to specifically not output those things verbatim specifically to thwart attempts to detect the presence of said works in training dataset (which would naturally reveal the model and its output being a derivative work).
Perhaps one could find some way of justifying that (people justified all sorts of stuff throughout history), but let it be something better than “the model is assumed to be a thinking human when it comes to IP abuse but unthinking tool when it comes to using it for personal benefit”.
[0] Of course, if you find me a single person on this planet capable of recalling 42% of any Harry Potter book, I’d be very impressed if I ever believed it.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#204That's a clickbait title. What they are actually saying: Given one correct quoted sentence, the model has 42% chance of predicting the next sentence correctly. So, assuming you start with the first sentence and tell it to keep going, it has a 0.42^n odds of staying on track, where n is the n-th sentence. It seems to me, that if they didn't keep correcting it over and over again with real quotes, it wouldn't even get…
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#205That's a clickbait title. What they are actually saying: Given one correct quoted sentence, the model has 42% chance of predicting the next sentence correctly. So, assuming you start with the first sentence and tell it to keep going, it has a 0.42^n odds of staying on track, where n is the n-th sentence. It seems to me, that if they didn't keep correcting it over and over again with real quotes, it wouldn't even get…
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#206Earlier quoted context omitted.
It seems like a pretty reasonable argument and easy enough to make. A human with a great memory could probably recreate some absurd % of Harry Potter after reading it, there are some very unusual minds out there. It is clear that if they read Harry Potter and being capable of reproducing it on demand as a party trick that would be fair use. So the LLM should also be fair use since it is using a mechanism similar enou…
> is clear that if they read Harry Potter and reproduce it on demand as a party trick that would be fair use. Actually no that could be copyright infringement. Badly signing a recent pop song in public also qualifies as copyright infringement. Public performances count as copying here.
For commercial purposes only. If someone sells a recreation of the Harry Potter book, it’s illegal regardless whether it was by memory, directly copying the book, or using an LLM. It’s the act of broadcasting it that’s infringing on copyright, not the content itself.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#207Earlier quoted context omitted.
It seems like a pretty reasonable argument and easy enough to make. A human with a great memory could probably recreate some absurd % of Harry Potter after reading it, there are some very unusual minds out there. It is clear that if they read Harry Potter and being capable of reproducing it on demand as a party trick that would be fair use. So the LLM should also be fair use since it is using a mechanism similar enou…
> is clear that if they read Harry Potter and reproduce it on demand as a party trick that would be fair use. Actually no that could be copyright infringement. Badly signing a recent pop song in public also qualifies as copyright infringement. Public performances count as copying here.
Although frankly, as has been pointed out many times, the law is also stupid in what it prohibits and that should be fixed first as a priority. Its done some terrible damage to our culture. My family used to be part of a community choir until it shut down basically for copyright reasons.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#208Earlier quoted context omitted.
Indeed but since when is a blatantly derived work only using 50% of a copyrighted work without permission a paragon of copyright compliance? Music artists get in trouble for using more than a sample without permission — imagine if they just used 45% of a whole song instead… I’m amazed AI companies haven’t been sued to oblivion yet. This utter stupidity only continues because we named a collection of matrices “Artific…
Music artists get in trouble for using more than a sample from other music artists without permission because their work is in direct competition with the work they're borrowing from. A ZIP file of a book is also in direct competition of the book, because you could open the ZIP file and read it instead of the book. A model that can take 50 tokens and give you a greater than 50% probability for the 50 next tokens 42%…
Under the hood they are 100% deterministic, modulo quantization and rounding errors.
So yes, it is very much possible to use LLMs as a lossy compressed archive for texts.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#209Do LLMs have any perception that Harry Potter is fiction or is it possible that they will give some magical advice based on fiction works that they have been trained with? edit: never mind, I’ll just ask ChatGPT