Live data from Hacker News

Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

understandingai.org

181–190 of 326 posts

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#181
post #137
post #133

Earlier quoted context omitted.

> Have you ever repeated a line from your favorite movie or TV show? Memorized a poem? Guess the rights holders better sue you for stealing their content by encoding it in your wetware neural network. I see this absolute non-argument regurgitated ad infinitum in every single discussion on this topic, and at this point I can't help but wonder: doesn't it say more about the person who says it than anything else? Do you…

Who said anything about freedom of speech? Nobody is claiming the LLM has free speech rights, which don't even apply to infringing copyright anyway. Freedom of speech doesn't give me the right to make copies of copyrighted works. The question is whether the model weights constitute of copy of the work. I contend that they do not, or they did, than so do the analogous weights (reinforced neural pathways) in your brain…

Making personal copies is generally permitted. If I were to distribute the neural pathways in my brain enabling others to reproduce copyrighted works verbatim, the owners of the copyrighted works would have a case against me.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#182
post #106

Earlier quoted context omitted.

Indeed but since when is a blatantly derived work only using 50% of a copyrighted work without permission a paragon of copyright compliance? Music artists get in trouble for using more than a sample without permission — imagine if they just used 45% of a whole song instead… I’m amazed AI companies haven’t been sued to oblivion yet. This utter stupidity only continues because we named a collection of matrices “Artific…

Music artists get in trouble for using more than a sample from other music artists without permission because their work is in direct competition with the work they're borrowing from. A ZIP file of a book is also in direct competition of the book, because you could open the ZIP file and read it instead of the book. A model that can take 50 tokens and give you a greater than 50% probability for the 50 next tokens 42%…

this is the first sensible argument in defense of AI models i read in this debate. thank you. this does make sense.

AI can reproduce individual sentences 42% of the time but it can't reproduce a summary.

the question however us, is that in the design if AI tools or us that a limitation of current models? what if future models get better at this and are able to produce summaries?

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#183
post #178

Earlier quoted context omitted.

I’ve yet to read an actual argument defending commercial LLM’s as fair use based on existing (edit:legal) criteria.

I'm yet to read an actual argument that it's not. Vibe-arguing "because corporations111" ain't it.

I’m looking for a link that does something like this but ends up supporting commercial LLM’s

https://copyrightalliance.org/faqs/what-is-fair-use/

The purpose and character of the use, including whether such use is of a commercial nature or is for non-profit educational purposes; (commercial least wiggle room) The nature of the copyrighted work; (fictional work least wiggle room) The amount and substantiality of the portion used in relation to the copyrighted work as a whole; (42% is considered a huge fraction of a book) and The effect of the use upon the potential market for or value of the copyrighted work. (Best argument as it’s minimal as a piece of entertainment. Not so as a cultural icon. Someone writing a book report or fan fiction may be less likely to buy a copy. )

Those aren’t the only factors, but I’m more interested in the counter argument here than trying to say they are copyright infringing.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#184
post #178

Earlier quoted context omitted.

It's not clear that it's incorrect.

I’ve yet to read an actual argument defending commercial LLM’s as fair use based on existing (edit:legal) criteria.

It seems like a pretty reasonable argument and easy enough to make. A human with a great memory could probably recreate some absurd % of Harry Potter after reading it, there are some very unusual minds out there. It is clear that if they read Harry Potter and being capable of reproducing it on demand as a party trick that would be fair use. So the LLM should also be fair use since it is using a mechanism similar enough to what humans do and what humans do is fine.

The LLMs I've used don't randomly start spouting Harry Potter quotes at me, they only bring it up if I ask. They aren't aiming to undermine copyright. And they aren't a very effective tool for it compared to the very well developed networks for pirating content. It seems to be a non-issue that will eventually be settled by the raw economic force that LLMs are bringing to bear on society in the same way that the movie industry ultimately lost the battle against torrents and had to compete with them.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#185

[flagged]

Please don't do this here. If you're going to use the word, use the word, but also, please don't use words like that about people here, no matter what you think of them. A comment like this breaks multiple guidelines:

https://news.ycombinator.com/newsguidelines.html

We detached this comment from https://news.ycombinator.com/item?id=44287156 and marked it off topic.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#186
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

> some massive new avenue to piracy

So it's fine as long as it's old piracy? How did you arrive to that conclusion?

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#187
post #63

Earlier quoted context omitted.

(Disclaimer: haven't read the original paper) It sounds like a ridiculous way to measure it. Producing 50-token excerpts absolutely doesn't translate to "recall X percent of Harry Potter" for me. (Edit: I read this article. Nothing burger if its interpretation of the original paper is correct.)

Their methodology seems reasonable to me. To clarify, they look at the probability a model will produce a verbatim 50-token excerpt given the preceding 50 tokens. They evaluate this for all sequences in the book using a sliding window of 10 characters (NB: not tokens). Sequences from Harry Potter have substantially higher probabilities of being reproduced than sequences from less well-known books. Whether this is "re…

> one of those tricky semantic arguments we have yet to settle when it comes to LLMs

Sure. But imagine this: In a hypothetical world where LLMs never ever exist, I tell you that I can recall 42 percent of the first Harry Potter book. What would you assume I can do?

It's definitely not "this guy can predict next 10 characters with 50% accuracy."

Of course the semantic of 'recall' isn't the point of this article. The point is that Harry Potter was in the training set. But I still think it's a nothing burger. It would be very weird to assume Llama was trained on copyright-free materials only. And afaik there isn't a legal precedent saying training on copyrighted materials is illegal.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#188
post #37

Earlier quoted context omitted.

But also we know for a fact that Meta trained their models on pirated books. So there's no need to invent a hare brained scheme of stitching together bits and pieces like that.

No, assuming that just because it was in the training data it must be memorized is hare brained. LLMs have limited capacity to memorize, under ~4 bits per parameter[1][2], and are trained on terabytes of data. It's physically impossible for them to memorize everything they're trained on. The model memorized chunks of Harry Potter not just because it was directly trained on the whole book, which the article also allud…

No, we know it because it was established in court from Meta internal communications.

https://www.theguardian.com/technology/2025/jan/10/mark-zuck...

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#189
post #109
post #104

Earlier quoted context omitted.

Of course it is. It's just a form of compression. If I train an autoencoder on an image, and distribute the weights, that would obviously be the same as distributing the content. Just because the content is commingled with lots of other content doesn't make it disappear. Besides, where did the sections of text from the input works that show up in the output text come from? Divine inspiration? God whispering to the ma…

[flagged]

Difference is if it's used commercially or not. Me singing my favourite song at karaoke is fine, but me recording that and releasing it on Spotify is not

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#190
post #184
post #178

Earlier quoted context omitted.

I’ve yet to read an actual argument defending commercial LLM’s as fair use based on existing (edit:legal) criteria.

It seems like a pretty reasonable argument and easy enough to make. A human with a great memory could probably recreate some absurd % of Harry Potter after reading it, there are some very unusual minds out there. It is clear that if they read Harry Potter and being capable of reproducing it on demand as a party trick that would be fair use. So the LLM should also be fair use since it is using a mechanism similar enou…

The difference might be the "human doing it as a party trick" vs "multi billion dollar corporation using it for profit".

Having said that I think the cat is very much out of the bag on this one and, personally, I think that LLMs should be allowed to be trained on whatever.

Post reply on HN