Live data from Hacker News

Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

understandingai.org

261–270 of 326 posts

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#261

Earlier quoted context omitted.

The main issue on an economical point of view is that copyright is not the framework we need for social justice and everyone florishing by enjoying pre-existing treasures of human heritage and fairly contributing back. There is no morale and justice ground to leverage on when the system is designed to create wealth bottleneck toward a few recipients. Harry Potter is a great piece of artistic work, and it's nice that…

Capitalism is allergic to second-order cybernetics. First-order systems drive outcomes. "Did it make money?" "Did it increase engagement?" "Did it scale?" These are tight, local feedback loops. They work because they close quickly and map directly to incentives. But they also hide a deeper danger: they optimize without questioning what optimization does to the world that contains it. Second-order cybernetics reason a…

I don't like many things about this post, its a bit snobbish and uses esoteric language in order to sound more intricate than it really is.

>Capitalism is not simply incapable of reflection. In fact, it's structured to ignore it. It has no native interest in what emerges from its aggregated behaviors unless those emergent properties threaten the throughput of capital itself. It isn't designed to ask, "What kind of society results from a thousand locally rational decisions?" It asks, "Is this change going to make more or less money?"

Capitalism and free market has lot of useful and emergent properties that occur not at the first order but second order.

> In the case of the global economic system, under capitalism, growth, accumulation and innovation can be considered emergent processes where not only does technological processes sustain growth, but growth becomes the source of further innovations in a recursive, self-expanding spiral. In this sense, the exponential trend of the growth curve reveals the presence of a long-term positive feedback among growth, accumulation, and innovation; and the emergence of new structures and institutions connected to the multi-scale process of growth

https://en.wikipedia.org/wiki/Emergence

In fact free market is an extremely good example of emergence or second order systems where each individual works selfishly but produces a second order effect of driving growth for everyone - something that is definitely preferable.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#262
post #22

It's important to note the way it was measured: > the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time. So 50 tokens is not really…

Even if it is recalling it 50 tokens at a time, the half of the book is in some sense in there, right?

Almost the entire book is in there. From the paper, if you give it a 100 token prompt, it will produce the next 50 tokens with more than 1% probability so that the produced tokens cover 91% of the book. And as the title says, it also produces next 50 tokens with more than 50% probability, so produced tokens cover 42% of the book. Bet it gets close to 100% as you reduce the probability.

Also they went through the book at 10 token strides. Like..a bit tortured way to reproduce the book (basically impossible to actually reproduce the book) but it shows that the content is in there.

Now whether this is derivative work, copyright violation or whatever is debatable. Probably gets similar numbers for a bunch of other books too. They should have done the Bible and probably get way higher numbers, but that won’t go viral.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#263
post #99
post #90

Earlier quoted context omitted.

You really don't see the difference between Google indexing the content of third parties and directly hosting/distributing the content itself?

Hosting model weights is not hosting / distributing the content.

I would be inclined to agree except apparently 42% of the first Harry Potter book is encoded in the model weights...

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#264
post #22

It's important to note the way it was measured: > the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time. So 50 tokens is not really…

[dead]

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#265
post #22

It's important to note the way it was measured: > the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time. So 50 tokens is not really…

An LLM is not a database. There is no significant amount of information in a model that can be accessed 100% of the time. This is because it's a mystery to the user what collection of tokens will lead to a specific output. To get a predictable result from an LLM 50% of the time is very significant.

This doesn't tell us for certain whether or not the model was trained on a full copy of the book. It's possible that 50-token long passages from 42% of the book were, incidentally, quoted verbatim in various parts of the training data. Considering the popularity of both the book itself, and derivative fan-fiction, I would not be surprised. I would be less surprised to learn that it was indeed trained on a full copy of the book, if not several.

The more meaningful point here is that the ability to reproduce half a book is the same sort of overt derivative work that is definitely considered copyright infringement in other circumstances. A lossy copy is still a copy. If we are to hold LLMs to the same standard as other content, this isn't very easy to defend.

Personally, I see this as a good opportunity to reevaluate copyright on the whole. I think we would be better off without it.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#266
post #215
post #178

Earlier quoted context omitted.

I’ve yet to read an actual argument defending commercial LLM’s as fair use based on existing (edit:legal) criteria.

Based upon legal decisions in the past there is a clear argument that the distinction for fair use is whether a work is substantially different to another. You are allowed to write a book containg information you learned about from another book. There is threshold in academia regarding plagiarism that stands apart from the legal standing. The measure that was used in Gyles v Wilcox was if the new work could substitut…

Training itself involves making infringing copies of protected works. Whether or not inference produces copyrighted material is almost beside the point.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#267

Earlier quoted context omitted.

Capitalism is allergic to second-order cybernetics. First-order systems drive outcomes. "Did it make money?" "Did it increase engagement?" "Did it scale?" These are tight, local feedback loops. They work because they close quickly and map directly to incentives. But they also hide a deeper danger: they optimize without questioning what optimization does to the world that contains it. Second-order cybernetics reason a…

I don't like many things about this post, its a bit snobbish and uses esoteric language in order to sound more intricate than it really is. >Capitalism is not simply incapable of reflection. In fact, it's structured to ignore it. It has no native interest in what emerges from its aggregated behaviors unless those emergent properties threaten the throughput of capital itself. It isn't designed to ask, "What kind of so…

Appreciate the engagement. But your reply mostly recenters a pro-capitalist narrative by redefining the products of "emergence" as inherently good. My argument isn't about stacking pros and cons and calculating the combined sum. It’s about a structural blind spot: capitalism systematically collapses higher-order questions about what kind of world were building into first-order value propositions like "growth," "utility," and "innovation."

That's the core problem. Capitalism resists second-order critique from within because it translates every possible value: justice, meaning, even critique itself, into terms it can price or optimize. Your response is a perfect example: you defend capitalism by listing its outputs, but that;s another first-order move. If you were engaging at the second-order level, you'd interrogate not what the system produces, but what it refuses to ask, and who gets to decide. That silence is precisely my point.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#268
post #184
post #178

Earlier quoted context omitted.

I’ve yet to read an actual argument defending commercial LLM’s as fair use based on existing (edit:legal) criteria.

It seems like a pretty reasonable argument and easy enough to make. A human with a great memory could probably recreate some absurd % of Harry Potter after reading it, there are some very unusual minds out there. It is clear that if they read Harry Potter and being capable of reproducing it on demand as a party trick that would be fair use. So the LLM should also be fair use since it is using a mechanism similar enou…

> It is clear that if they read Harry Potter and being capable of reproducing it on demand as a party trick that would be fair use.

Not fair use. No one would ever prosecute it as infringement but it's not fair use.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#269
post #22

It's important to note the way it was measured: > the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time. So 50 tokens is not really…

Suppose for simplicity that every sentence in the book is 50 tokens or shorter. According to the stated methodology, I could give the LLM sentence 1 and have 42% chance of getting sentence 2 recalled. Then I could give it sentence 2 and have 42% chance of getting sentence 3. Therefore, the LLM contains 42% of the book in some sense . I disagree this is "not really very much". If a person could do this you would undou…

Except it's not what happened, per the article. Instead, they walked down the logits, which is more like asking someone to give 10-20 best guesses for next word, and should one of them match the secret answer, telling them which one is it and asking them to go on with the next word. Seems like a substantially easier task, and most of information is coming from researchers making a choice at every step.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#270

Earlier quoted context omitted.

Is it Llama violating the "copyright" or is it the researcher pushing it to do so?

If you distribute a zip file of the book, are you violating copyright, or is it the person who unzips it?

You are.

Copyright is quite literally about the right to control the creation and distribution of copies.

The creation of the unzipped file is not treated as a separate copy so the recipient would not be violating copyright just by unzipping the file you provided.

Post reply on HN