Live data from Hacker News

Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

understandingai.org

241–250 of 326 posts

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#241
post #22

It's important to note the way it was measured: > the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time. So 50 tokens is not really…

Even if it is recalling it 50 tokens at a time, the half of the book is in some sense in there, right?

Not necessarily. Information is always spread between what we'd normally consider "storage medium" and "reader"; the degree to which that is is a controllable parameter.

Consider e.g.:

- Digital expansion of PI to sufficient decimal places contains both parts of the work and full work in full. The trick is you have to know where to find it - and it's that knowledge that's actually equivalent to the work itself.

- Any kind of compression that uses a dictionary that's separate from the compressed artifact, shifts some of the information into a dictionary file, or if it's a common dictionary, into compressor/decompressor itself.

In the case from the study, the experimenter actually has to supply most of the information required to pull Harry Potter out of the model - they need to make specific prompts with quotes from the book, and then observe which logits correspond to the actual continuation of those quotes. The experimenter is doing information-loaded selection multiple times: at prompting, and at identifying logits. This by itself doesn't really prove the model memorized the book, only just that it saw fragments from it - in cases those fragments are book-specific (e.g. using proper names from the HP world) instead of generic English sentences.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#242
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

You're attacking a strawman. Nobody's claiming LLMs are a new piracy vector or that people will use ChatGPT, Llama or Claude instead of buying Harry Potter.

The issue here is that tech companies systematically copied millions of copyrighted works to build commercial products worth billions, without reembursing the people who made their products possible in the first place. The research shows Llama literally memorized 42% of Harry Potter - not simply "learned from it," but can reproduce it verbatim. That's 1) not transformative and 2) clear evidence of copyright infringement.

By your logic, the existence of torrents would make it perfectly acceptable for someone to download pirated movies and charge people to stream them. "Piracy already exists" isn't a defense, and it especially shouldn't be for companies worth billions. But you bet your ass that if I built a commercial Netflix competitor built on top of systematic copyright violations, I'd be sued into the dirt faster than I can say "billion dollar valuation".

Aaron Swartz faced 35 years in prison and ultimately took his own life over downloading academic papers that were largely publicly funded. He wasn't selling them, he wasn't building a commercial product worth billions of dollars - he was trying to make knowledge accessible.

Meanwhile, these AI companies like Meta systematically ingested copyrighted works at an industrial scale to build products worth billions. Why does an individual face life-destroying prosecution for far less, while trillion dollar companies get to negotiate in civil court after building empires on others' works? And why are you defending them?

Edit:

And for what it's worth, I'm far from a copyright maximalist. I've long believed that copyright terms - especially decades after creators' deaths - have become excessive. But whatever your stance on copyright ultimately is, the rules should apply equally to individuals like Aaron and multi-billion dollar corporations.

You cannot seriously use the fact that individuals may pirate a book (which is illegal) as an ethical or legal defense for corporations doing the same thing at an industrial scale for profit.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#243
post #184

Earlier quoted context omitted.

It seems like a pretty reasonable argument and easy enough to make. A human with a great memory could probably recreate some absurd % of Harry Potter after reading it, there are some very unusual minds out there. It is clear that if they read Harry Potter and being capable of reproducing it on demand as a party trick that would be fair use. So the LLM should also be fair use since it is using a mechanism similar enou…

> A human with a great memory This kind of argument keeps popping up usually to justify why training LLMs on protected material is fair, and why their output is fair. It's always used in a super selective way, never accounting for confounding factors, just because superficially it sort of supports that idea. Exceptional humans are exceptional, rare. When they learn, or create something new based on prior knowledge, o…

Nobody in real life thinks humans and machines are the same thing and actually believes they should have the same legal status. The A.I. enthusiast would not support the legality of shooting them when no longer useful the way a company would shred an old hard drive.

This supposed failure to see the difference between the human mind and a machine whenever someone brings up copyright is peformative and disingenuous.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#244

Well, so can a nontrivial number of people. It's Harry Potter we're talking about - it's up there with The Bible in popularity ranking. I'm gonna bet that Llama 3.1 can recall a significant portion of Pride and Prejudice too. With examples of this magnitude, it's normal and entirely expected this can happen - as it does with people[0] - the only thing this is really telling us is that the model doesn't understand its…

Agree completely. When I read the Gemma 3 paper (https://arxiv.org/html/2503.19786v1) and saw an entire section dedicated to measuring and reducing the memorization rate I was annoyed. How does this benefit end users at all?

I want the language model I'm using to have knowledge of cultural artifacts. Gemma 3 27B was useless at a question related to grouping Berserk characters by potential baldurs gate 3 classes; Claude did fine. The methods used to reduce memorisation rate probably also deteriorate performance in some other ways that don't show up on benchmarks.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#245
post #101

Earlier quoted context omitted.

> I can record myself reciting the full Harry Potter book then distribute it on YouTube Not legally you can't. Both of your examples are copyright violations

Recording yourself is not a violation, only publishing on Youtube. Content generated with LLMs are not a violation. Publishing the content you generated might be.

Generating the content for the user is the distribution regardless of what the user does with it

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#246
post #22

It's important to note the way it was measured: > the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time. So 50 tokens is not really…

> 50 tokens is not really very much Yes! And also llama3.1’s tokens are different from Qwen and llama1 tokens. That’s the first model where meta started to use very large vocab_size.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#247
post #218
post #215

Earlier quoted context omitted.

Based upon legal decisions in the past there is a clear argument that the distinction for fair use is whether a work is substantially different to another. You are allowed to write a book containg information you learned about from another book. There is threshold in academia regarding plagiarism that stands apart from the legal standing. The measure that was used in Gyles v Wilcox was if the new work could substitut…

I think you may have something with that line of reasoning. The threshold for transformative for fictional works is fairly high unfortunately. Fan fiction and reasonably distinct works with excessive inspiration are both copyright infringing. https://en.wikipedia.org/wiki/Tanya_Grotter > Models themselves are very clearly transformative. A near word for word copy of large sections of a work seems nowhere near that th…

Models are not word for word copies of large sections of text. They are capable of emitting that text though.

It would be interesting to look at what legal precidents were set regarding mp3s or other encodings. Is the encoding itself an infringement, or is it the decoding, or is it the distribution of a decodable form of a work.

There is also the distinction with a lossy encoding that encodes a single work. There is clarity when the encoded form serves no other purpose other than to be decoded into a given work. When the encoding acts as a bulk archive, does the responsibility shift to those who choose what to extract from the archive?

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#248
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

Everything you mentioned can simply be deleted. You can't really delete this from the "brain" of the LLM if a court orders you to do so, you have to re-train the LLM, which is costly. That's the problem I see.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#250
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

Problem is that it copies much more work than just harry potter, including yours if you ever shared it (even under copy-left license) and makes money off it.
Post reply on HN