Live data from Hacker News

Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

understandingai.org

41–50 of 326 posts

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#41

I really wish we could get rid of copyright. It's going to hold us back long term.

We cannot get ride of it without finding a way to pay the creators that generate copyrighted works.

I’m personally more in favor of significantly reducing the length of the copy right. I think 20-30 years is an interesting range. Artist get roughly a career length of time to profit off their creations, but there is much less incentive for major corporations to buy and horde IP.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#42
post #17

Earlier quoted context omitted.

That's got nothing to do with it. It's all about copyright. Can it reproduce its training data verbatim? If so, Meta is in hot water.

I read harry potter, and you ask me about a page, and I can recite it verbatim, did I just commit copyright infringement?

This is an extremely common strawman argument. We're not discussing human memory.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#43
post #17

Earlier quoted context omitted.

That's got nothing to do with it. It's all about copyright. Can it reproduce its training data verbatim? If so, Meta is in hot water.

I read harry potter, and you ask me about a page, and I can recite it verbatim, did I just commit copyright infringement?

I pay for a service. The service recites a novel to me. The service would need permission to do this or it is copyright infringement.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#44
post #22

It's important to note the way it was measured: > the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time. So 50 tokens is not really…

Fair use is a four part test, and the amount if copying is only one of the four parts.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#45

As I've said several times, the corpus is key: LLMs thus far "read" most anything, but should instead have well-curated corpora. "Garbage In, Garbage Out!(GIGO)" is the saying. While the Harry Potter series may be fun reading, it doesn't provide information about anything that isn't better covered elsewhere. Leave Harry Potter for a different "Harry Potter LLM". Train scientific LLMs to the level of a good early 20th…

> While the Harry Potter series may be fun reading, it doesn't provide information about anything that isn't better covered elsewhere

To address this point, and not other concerns: the benefits would be (1) pop culture knowledge and (2) having a variety of styles of edited/reasonably good-quality prose.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#46
post #8

From a quick web search I can find that there are book review sites that allow users to enter and rate verbatim "quotes" from books. This one [1] contains ~2000 [2] portions of a sentence, a paragraph or several paragraphs of Harry Potter and the Sorcerer's Stone. Could it be plausible that an LLM had ingested parts of the book via scrapping web pages like this and not the full copyrighted book and get results simila…

Sure, why not? lol https://www.reddit.com/r/DataHoarder/comments/1entowq/i_made... https://github.com/shloop/google-book-scraper The fact that Meta torrented Books3 and other datasets seems to be by self-admission by Meta employees who performed the work and/or oversaw those who themselves did the work, so that is not really under dispute or ambiguous. https://torrentfreak.com/meta-admits-use-of-pirated-book-dat...

Books3 was used in Llama1. We don't know if they used it later on.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#47
post #17

Earlier quoted context omitted.

That's got nothing to do with it. It's all about copyright. Can it reproduce its training data verbatim? If so, Meta is in hot water.

I read harry potter, and you ask me about a page, and I can recite it verbatim, did I just commit copyright infringement?

Are you selling your ability to recite stuff? Then certainly.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#48
post #36
post #22

It's important to note the way it was measured: > the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time. So 50 tokens is not really…

This sounds almost like "Works every time (50% of the time)."

Except the odds of it happening even 50% of the time is less likely than winning the lottery multiple times. All while illegally ingesting copywrite material without (and presumably against the wishes of) the consent of the copywrite holder.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#49
post #25

Earlier quoted context omitted.

> the physical encoding which definitely exists in my brain is a copyright violation First of all, we don't really know how the brain works. I get that you're being a snarky physicalist, but there's plenty of substance dualists, panpsychsts, etc. out there. So, some might say, this is a reductive description of what happens in our brains. Second of all, yes, if you tried to publish Harry Potter ( even if it was from…

Right but the physical encoding already exists in my brain or how can I reproduce it in the first place? We may not know how the encoding works but we do know that an encoding exists because a decoding is possible. My question is… is that in itself a violation of copyright? If not then as long as LLMs don’t make a publication it shouldn’t be a copyright violation right? Because we don’t understand how it’s encoded in…

I don’t think the lawyers are going to buy arguments that compare LLMs with human biology like this.
Post reply on HN