I'm surprised no one in the comments has mentioned overfitting. Perhaps this is too obvious but I think of it as a very clear bug in a model if it asserts something to be true because it has heard it once. I realize that training a model is not easy, but this is something that should've been caught before it was released. Either QA is sleeping on the job or they have intentionally released a model with serious flaws…
Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
171–180 of 326 posts
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#172Earlier quoted context omitted.
> Freedom of speech doesn't give me the right to make copies of copyrighted works. No, but it gives you the right to quote a line from a movie or TV show without being charged with copyright infringement. You argued that an LLM deserves that same right, even if you didn't realize it. > than so do the analogous weights (reinforced neural pathways) in your brain Did your brain consume millions of copyrighted books in o…
Millions? No, but my brain certainly consumed thousands of books, movies, TV shows, pieces of music, artworks, and other copyrighted material. Where is the cutoff? Can I only consume 999,999 copyrighted works before I'm not longer allowed to remember something without infringing copyright? My brain definitely would not exist in its current form without consuming that material. It would exist in some form, but it woul…
this is literally why i don't like to work on proprietary code. because when i need to create a similar solution for someone else i have to go out of my way to make sure i do it differently. people have been sued over this.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#173As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…
People aren't buying Harry Potter action figures as a subtitute for buying the book either, but copyright protects creators from other people swooping in and using their work in other mediums. There is obviously a huge market demand for high quality data for training LLMs, Meta just spent 15 billion on a data labeling company. Companies training LLMs on copyrighted material without permission are doing that as a subs…
Dropping the novels into a machine‑learning corpus is a fundamentally different act. The text is not being resold, and the resulting model is not advertised as “official Harry Potter.” The books are just statistical nutrition. One ingredient among millions. Much like a human writer who reads widely before producing new work. No consumer is choosing between “Rowling’s novel” and “the tokens her novel contributed to an LLM,” so there’s no comparable displacement of demand.
In economic terms, the merch market is rivalrous and zero‑sum; the training market is non‑rivalrous and produces no direct substitute good. That asymmetry is why copyright doctrine (and fair‑use case law) treats toy knock‑offs and corpus building very differently.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#174Earlier quoted context omitted.
That can easily be written off as fair use. No, it really couldn't. In fact, it's very persuasive evidence that Llama is straight up violating copyright. It would be one thing to be able to "predict" a paragraph or two. It's another thing entirely to be able to predict 42% of a book that is several hundred pages long.
Is it Llama violating the "copyright" or is it the researcher pushing it to do so?
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#175Earlier quoted context omitted.
> let's not pretend that an LLM that autocompletes a couple lines from harry potter with 50% accuracy is some massive new avenue to piracy No one is claiming this. The corporations developing LLMs are doing so by sampling media without their owners' permission and arguing this is protected by US fair use laws, which is incorrect - as the late AI researcher Suchir Balaji explained in this other article: https://suchir…
It's not clear that it's incorrect.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#176Earlier quoted context omitted.
> let's not pretend that an LLM that autocompletes a couple lines from harry potter with 50% accuracy is some massive new avenue to piracy No one is claiming this. The corporations developing LLMs are doing so by sampling media without their owners' permission and arguing this is protected by US fair use laws, which is incorrect - as the late AI researcher Suchir Balaji explained in this other article: https://suchir…
Yeah, that's literally the title of the article,and the premise of the first paragraph.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#177I'm gonna bet that Llama 3.1 can recall a significant portion of Pride and Prejudice too.
With examples of this magnitude, it's normal and entirely expected this can happen - as it does with people[0] - the only thing this is really telling us is that the model doesn't understand its position in the society well enough to know to shut up; that obliging the request is going to land it, or its owners, into trouble.
In some way, it's actually perverted.
EDIT: it's even worse than that. What the research seems to be measuring is that the models recognize sentence-sized pieces of the book as likely continuations of an earlier sentence-sized piece. Not whether it'll reproduce that text when used straightforwardly - just whether there's an indication it recognizes the token patterns as likely.
By that standard, I bet there's over a billion people right now who could do that to 42% of first Harry Potter book. By that standard, I too memorized the Bible end-to-end, as had most people alive today, whether or not they're Christian; works this popular bleed through into common language usage patterns.
--
[0] - Even more so when you relax your criteria to accept occasional misspell or paraphrase - then each of us likely know someone who could piece together a chunk of HP book from memory.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#178Earlier quoted context omitted.
> let's not pretend that an LLM that autocompletes a couple lines from harry potter with 50% accuracy is some massive new avenue to piracy No one is claiming this. The corporations developing LLMs are doing so by sampling media without their owners' permission and arguing this is protected by US fair use laws, which is incorrect - as the late AI researcher Suchir Balaji explained in this other article: https://suchir…
It's not clear that it's incorrect.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#179Earlier quoted context omitted.
Of course it is. It's just a form of compression. If I train an autoencoder on an image, and distribute the weights, that would obviously be the same as distributing the content. Just because the content is commingled with lots of other content doesn't make it disappear. Besides, where did the sections of text from the input works that show up in the output text come from? Divine inspiration? God whispering to the ma…
[flagged]
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#180Earlier quoted context omitted.
It's not clear that it's incorrect.
I’ve yet to read an actual argument defending commercial LLM’s as fair use based on existing (edit:legal) criteria.
Vibe-arguing "because corporations111" ain't it.