Live data from Hacker News

Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

understandingai.org

171–180 of 326 posts

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#171

I'm surprised no one in the comments has mentioned overfitting. Perhaps this is too obvious but I think of it as a very clear bug in a model if it asserts something to be true because it has heard it once. I realize that training a model is not easy, but this is something that should've been caught before it was released. Either QA is sleeping on the job or they have intentionally released a model with serious flaws…

Overfitting makes for more human-like output (because it's repeating words written by a human). Out of all possible failure states of a model, overfitting is probably what you want out of an LLM, as long as it's not overfitted enough to lose lawsuits.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#172
post #146
post #141

Earlier quoted context omitted.

> Freedom of speech doesn't give me the right to make copies of copyrighted works. No, but it gives you the right to quote a line from a movie or TV show without being charged with copyright infringement. You argued that an LLM deserves that same right, even if you didn't realize it. > than so do the analogous weights (reinforced neural pathways) in your brain Did your brain consume millions of copyrighted books in o…

Millions? No, but my brain certainly consumed thousands of books, movies, TV shows, pieces of music, artworks, and other copyrighted material. Where is the cutoff? Can I only consume 999,999 copyrighted works before I'm not longer allowed to remember something without infringing copyright? My brain definitely would not exist in its current form without consuming that material. It would exist in some form, but it woul…

i can remember and i can quote, but if i quote to much i violate the copyright.

this is literally why i don't like to work on proprietary code. because when i need to create a similar solution for someone else i have to go out of my way to make sure i do it differently. people have been sued over this.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#173
post #143
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

People aren't buying Harry Potter action figures as a subtitute for buying the book either, but copyright protects creators from other people swooping in and using their work in other mediums. There is obviously a huge market demand for high quality data for training LLMs, Meta just spent 15 billion on a data labeling company. Companies training LLMs on copyrighted material without permission are doing that as a subs…

Harry Potter action figures trade almost entirely on J. K. Rowling’s expressive choices. Every unlicensed toy competes head‑to‑head with the licensed one and slices off a share of a finite pot of fandom spending. Copyright law treats that as classic market substitution and rightfully lets the author police it.

Dropping the novels into a machine‑learning corpus is a fundamentally different act. The text is not being resold, and the resulting model is not advertised as “official Harry Potter.” The books are just statistical nutrition. One ingredient among millions. Much like a human writer who reads widely before producing new work. No consumer is choosing between “Rowling’s novel” and “the tokens her novel contributed to an LLM,” so there’s no comparable displacement of demand.

In economic terms, the merch market is rivalrous and zero‑sum; the training market is non‑rivalrous and produces no direct substitute good. That asymmetry is why copyright doctrine (and fair‑use case law) treats toy knock‑offs and corpus building very differently.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#174

Earlier quoted context omitted.

That can easily be written off as fair use. No, it really couldn't. In fact, it's very persuasive evidence that Llama is straight up violating copyright. It would be one thing to be able to "predict" a paragraph or two. It's another thing entirely to be able to predict 42% of a book that is several hundred pages long.

Is it Llama violating the "copyright" or is it the researcher pushing it to do so?

If you distribute a zip file of the book, are you violating copyright, or is it the person who unzips it?

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#175
post #156

Earlier quoted context omitted.

> let's not pretend that an LLM that autocompletes a couple lines from harry potter with 50% accuracy is some massive new avenue to piracy No one is claiming this. The corporations developing LLMs are doing so by sampling media without their owners' permission and arguing this is protected by US fair use laws, which is incorrect - as the late AI researcher Suchir Balaji explained in this other article: https://suchir…

It's not clear that it's incorrect.

[deleted]

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#176
post #156

Earlier quoted context omitted.

> let's not pretend that an LLM that autocompletes a couple lines from harry potter with 50% accuracy is some massive new avenue to piracy No one is claiming this. The corporations developing LLMs are doing so by sampling media without their owners' permission and arguing this is protected by US fair use laws, which is incorrect - as the late AI researcher Suchir Balaji explained in this other article: https://suchir…

Yeah, that's literally the title of the article,and the premise of the first paragraph.

The first paragraph isn’t arguing that this copying will lead to piracy. It’s referring to court cases where people are trying to argue LLM’s themselves are copyright infringing.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#177
Well, so can a nontrivial number of people. It's Harry Potter we're talking about - it's up there with The Bible in popularity ranking.

I'm gonna bet that Llama 3.1 can recall a significant portion of Pride and Prejudice too.

With examples of this magnitude, it's normal and entirely expected this can happen - as it does with people[0] - the only thing this is really telling us is that the model doesn't understand its position in the society well enough to know to shut up; that obliging the request is going to land it, or its owners, into trouble.

In some way, it's actually perverted.

EDIT: it's even worse than that. What the research seems to be measuring is that the models recognize sentence-sized pieces of the book as likely continuations of an earlier sentence-sized piece. Not whether it'll reproduce that text when used straightforwardly - just whether there's an indication it recognizes the token patterns as likely.

By that standard, I bet there's over a billion people right now who could do that to 42% of first Harry Potter book. By that standard, I too memorized the Bible end-to-end, as had most people alive today, whether or not they're Christian; works this popular bleed through into common language usage patterns.

--

[0] - Even more so when you relax your criteria to accept occasional misspell or paraphrase - then each of us likely know someone who could piece together a chunk of HP book from memory.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#178
post #156

Earlier quoted context omitted.

> let's not pretend that an LLM that autocompletes a couple lines from harry potter with 50% accuracy is some massive new avenue to piracy No one is claiming this. The corporations developing LLMs are doing so by sampling media without their owners' permission and arguing this is protected by US fair use laws, which is incorrect - as the late AI researcher Suchir Balaji explained in this other article: https://suchir…

It's not clear that it's incorrect.

I’ve yet to read an actual argument defending commercial LLM’s as fair use based on existing (edit:legal) criteria.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#179
post #109
post #104

Earlier quoted context omitted.

Of course it is. It's just a form of compression. If I train an autoencoder on an image, and distribute the weights, that would obviously be the same as distributing the content. Just because the content is commingled with lots of other content doesn't make it disappear. Besides, where did the sections of text from the input works that show up in the output text come from? Divine inspiration? God whispering to the ma…

[flagged]

Repeating half of the book verbatim is not nearly the same as repeating a line.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#180
post #178

Earlier quoted context omitted.

It's not clear that it's incorrect.

I’ve yet to read an actual argument defending commercial LLM’s as fair use based on existing (edit:legal) criteria.

I'm yet to read an actual argument that it's not.

Vibe-arguing "because corporations111" ain't it.

Post reply on HN