Live data from Hacker News

Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

understandingai.org

251–260 of 326 posts

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#251
post #244

Well, so can a nontrivial number of people. It's Harry Potter we're talking about - it's up there with The Bible in popularity ranking. I'm gonna bet that Llama 3.1 can recall a significant portion of Pride and Prejudice too. With examples of this magnitude, it's normal and entirely expected this can happen - as it does with people[0] - the only thing this is really telling us is that the model doesn't understand its…

Agree completely. When I read the Gemma 3 paper ( https://arxiv.org/html/2503.19786v1 ) and saw an entire section dedicated to measuring and reducing the memorization rate I was annoyed. How does this benefit end users at all? I want the language model I'm using to have knowledge of cultural artifacts. Gemma 3 27B was useless at a question related to grouping Berserk characters by potential baldurs gate 3 classes; Cl…

> When I read the Gemma 3 paper (https://arxiv.org/html/2503.19786v1) and saw an entire section dedicated to measuring and reducing the memorization rate I was annoyed. How does this benefit end users at all?

It benefits users because memorisation is a waste of parameters that would be more useful if they were instead learning rules and generalisations.

For short snippets, common idioms and quotations that people recognise, exact quotes can be worth memorising; but the longer the quotations get, the less often it is important to be word-for-word exact — even for just a few paragraphs, I think most people only ever do oaths, anthems, songs they really like, and possibly a few hobbies.

If you want an exact quote, use (or tell the AI to use) a search engine.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#252
post #22

It's important to note the way it was measured: > the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time. So 50 tokens is not really…

Suppose for simplicity that every sentence in the book is 50 tokens or shorter.

According to the stated methodology, I could give the LLM sentence 1 and have 42% chance of getting sentence 2 recalled. Then I could give it sentence 2 and have 42% chance of getting sentence 3. Therefore, the LLM contains 42% of the book in some sense.

I disagree this is "not really very much". If a person could do this you would undoubtedly conclude that the person read the book.

In fact the number 42% even understates the severity of the matter. Superficially it makes it sound that the LLM only contains less than half of the book. In reality the process I described applies to 100% of the sentences. Additionally I'm guessing that the 58% times where the 50 tokens arent recalled correctly, the outputted token probably have the same meaning as the correct one.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#253

Well, so can a nontrivial number of people. It's Harry Potter we're talking about - it's up there with The Bible in popularity ranking. I'm gonna bet that Llama 3.1 can recall a significant portion of Pride and Prejudice too. With examples of this magnitude, it's normal and entirely expected this can happen - as it does with people[0] - the only thing this is really telling us is that the model doesn't understand its…

I keep waiting for the day when software stops being compared to a human person (a being with agency, free will, consciousness, and human rights of its own) for the purposes of justifying IP law circumvention. Yes, there is no problem when a person reads some book and recalls pieces[0] of it in a suitable context. How would that in any way address when certain people create and distribute commercial software, providi…

> I keep waiting for the day when software stops being compared to a human person (a being with agency, free will, consciousness, and human rights of its own) for the purposes of justifying IP law circumvention.

I mean, "agency" is a goal of some AI; "free will" is incoherent*; the word "consciousness" has about 40 different definitions, some of which are so broad they include thermostats and others so narrow that it's provably impossible for anything (including humans) to have it; and "human rights" are a purely legal concept.

> What’s even worse, is that imaginably they train (or would train) the models to specifically not output those things verbatim specifically to thwart attempts to detect the presence of said works in training dataset (which would naturally reveal the model and its output being a derivative work).

Some of the makers certainly do as you say; but also, the more verbatim quotations a model can produce, the more computational effort that model needs to spend to get the far more useful general purpose results.

* I'm not a fan of Aleister Crowley, but I think he was right to say that there's only one thing you can actually do that's truly your own will and not merely you allowing others to influence you: https://en.wikipedia.org/wiki/True_Will

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#255

Earlier quoted context omitted.

> A human with a great memory This kind of argument keeps popping up usually to justify why training LLMs on protected material is fair, and why their output is fair. It's always used in a super selective way, never accounting for confounding factors, just because superficially it sort of supports that idea. Exceptional humans are exceptional, rare. When they learn, or create something new based on prior knowledge, o…

Nobody in real life thinks humans and machines are the same thing and actually believes they should have the same legal status. The A.I. enthusiast would not support the legality of shooting them when no longer useful the way a company would shred an old hard drive. This supposed failure to see the difference between the human mind and a machine whenever someone brings up copyright is peformative and disingenuous.

> Nobody in real life thinks humans and machines are the same thing

Maybe you've been following a different conversation, or jumping to conclusions is just more convenient. This isn't about "legal status of AI" but about laws written having in mind only the capabilities of humans, at a time when systems as powerful as today's were unthinkable. Obviously the same laws have to set different limits for humans and machines.

There's no law limiting a human's top (running) speed but you have speed limits for cars. Maybe you're legally allowed to own a semi-automatic weapon but not an automatic one. This is the ELI5 for why when legislating, capabilities make all the difference. Obviously a rifle should not have the same legal status or be the same thing as a human, just in case my point is still lost on you.

Literally every single discussion on this LLM training/output topic, this one included, eventually has a number of people basing their argument on "but humans are allowed to do it", completely ignoring that humans can only do it in a much, much more limited way.

> is peformative and disingenuous

That's an extremely uncharitable and aggressive take, especially after not bothering to understand at all what I said.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#256
post #247
post #218

Earlier quoted context omitted.

I think you may have something with that line of reasoning. The threshold for transformative for fictional works is fairly high unfortunately. Fan fiction and reasonably distinct works with excessive inspiration are both copyright infringing. https://en.wikipedia.org/wiki/Tanya_Grotter > Models themselves are very clearly transformative. A near word for word copy of large sections of a work seems nowhere near that th…

Models are not word for word copies of large sections of text. They are capable of emitting that text though. It would be interesting to look at what legal precidents were set regarding mp3s or other encodings. Is the encoding itself an infringement, or is it the decoding, or is it the distribution of a decodable form of a work. There is also the distinction with a lossy encoding that encodes a single work. There is…

[deleted]

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#257

Earlier quoted context omitted.

Capitalism is allergic to second-order cybernetics. First-order systems drive outcomes. "Did it make money?" "Did it increase engagement?" "Did it scale?" These are tight, local feedback loops. They work because they close quickly and map directly to incentives. But they also hide a deeper danger: they optimize without questioning what optimization does to the world that contains it. Second-order cybernetics reason a…

Copyright doesn't "produce a cultural hellscape." That's just nonsense. Capitalism does because it has editorial control over narratives and their marketing and distribution. Those are completely different phenomena. Removing copyright will not suddenly open the floodgates of creativity because anyone can already create anything. But - and this is the key point - most work is me-too derivative anyway . See for exampl…

Everything is derivative. This boundary you are defending between originality and slop is extremely subjective at best. What harm is slop anyway? If originality is so objectively valuable, then why should its value be systemically enforced?

At the intersection of capitalism and copyright, I see a serious problem. Collaboration is encapsulated by competition. Because simple derivative work is illegal, all collaboration must be done in teams. Copyright defines every work of art as an island, whose value is not the art itself, but the moat that surrounds it. It should be no surprise that giant anticompetitive corporations reflect this structure. The core value of copyright is not creativity: it's rent-seeking.

Without copyright, we could collaborate freely. Our work would not be required to compete at all! Instead of victory over others' work, our goal could be success!

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#258
post #167

Earlier quoted context omitted.

Capitalism is allergic to second-order cybernetics. First-order systems drive outcomes. "Did it make money?" "Did it increase engagement?" "Did it scale?" These are tight, local feedback loops. They work because they close quickly and map directly to incentives. But they also hide a deeper danger: they optimize without questioning what optimization does to the world that contains it. Second-order cybernetics reason a…

and as a consequence the fight of AI vs copyright is one of two capitalists fighting each other. it's not about liberating copyright but about shuffling profits around. regardless of who wins that fight society loses. it conjures up pictures of two dragons fighting each other instead of attacking us, but make no mistake they are only fighting for the right to attack us. whoever wins is coming for us afterwards

The AI companies want two things:

1. Strong copyright to prevent competition from undercutting their related businesses.

2. Exclusive rights to totally ignore the copyright of everyone that made the content they use to train models.

I personally would much prefer we take the opportunity to abolish copyright entirely: for everyone, not just a handful of corporations. If derivative work is so valuable to our society (I believe it is), then I should be free to derive NVIDIA's GPU drivers without permission.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#259
post #247
post #218

Earlier quoted context omitted.

I think you may have something with that line of reasoning. The threshold for transformative for fictional works is fairly high unfortunately. Fan fiction and reasonably distinct works with excessive inspiration are both copyright infringing. https://en.wikipedia.org/wiki/Tanya_Grotter > Models themselves are very clearly transformative. A near word for word copy of large sections of a work seems nowhere near that th…

Models are not word for word copies of large sections of text. They are capable of emitting that text though. It would be interesting to look at what legal precidents were set regarding mp3s or other encodings. Is the encoding itself an infringement, or is it the decoding, or is it the distribution of a decodable form of a work. There is also the distinction with a lossy encoding that encodes a single work. There is…

> Is the encoding itself an infringement

Barring a fair use exception, yes.

From what I’ve read MP3’s get the same treatment as cassette tapes which were also lossy. It’s 1:1 digital copies that represented some novelty, but that rarely matters.

I’m hesitant to comment of the rest of that. The ultimate question isn’t if some difference exists but why that difference matters.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#260

Earlier quoted context omitted.

Nobody in real life thinks humans and machines are the same thing and actually believes they should have the same legal status. The A.I. enthusiast would not support the legality of shooting them when no longer useful the way a company would shred an old hard drive. This supposed failure to see the difference between the human mind and a machine whenever someone brings up copyright is peformative and disingenuous.

> Nobody in real life thinks humans and machines are the same thing Maybe you've been following a different conversation, or jumping to conclusions is just more convenient. This isn't about "legal status of AI" but about laws written having in mind only the capabilities of humans, at a time when systems as powerful as today's were unthinkable. Obviously the same laws have to set different limits for humans and machin…

>That's an extremely uncharitable and aggressive take, especially after not bothering to understand at all what I said.

To be clear, my intent wasn't to say you were the one being performative and disingenuous. I was referring to the sort of person you were debating against, the one who thinks every legal issue involving A.I. can be settled by typing "humans are allowed to do it."

Since I replied to you, I can see how what I wrote was confusing. My apologies.

The parent you replied to claimed LLMs are using "mechanism similar enough to what humans do and what humans do is fine."

Parent probably doesn't want his or her brain shredded like an old hard drive despite claiming similar mechanisms whenever it is convinient.

I'm arguing nobody actually believes there are "similar mechanisms" between machines and humans in their revealed preferences in day to day life.

>There's no law limiting a human's top (running) speed but you have speed limits for cars. Maybe you're legally allowed to own a semi-automatic weapon but not an automatic one.

I don't believe this analogy works. If we're talking about transmitting the text of Harry Potter, I believe it would already be illegal for a single human to type it on demand as a service.

If we are talking about remembering the text of Harry Potter but not reciting it on demand, that's not illegal for a human because copyright doesn't govern human memories.

I don't see what copyright law you think needs updating.

Post reply on HN