Live data from Hacker News

Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

understandingai.org

11–20 of 326 posts

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#11
post #8

From a quick web search I can find that there are book review sites that allow users to enter and rate verbatim "quotes" from books. This one [1] contains ~2000 [2] portions of a sentence, a paragraph or several paragraphs of Harry Potter and the Sorcerer's Stone. Could it be plausible that an LLM had ingested parts of the book via scrapping web pages like this and not the full copyrighted book and get results simila…

Sure, why not? lol

https://www.reddit.com/r/DataHoarder/comments/1entowq/i_made...

https://github.com/shloop/google-book-scraper

The fact that Meta torrented Books3 and other datasets seems to be by self-admission by Meta employees who performed the work and/or oversaw those who themselves did the work, so that is not really under dispute or ambiguous.

https://torrentfreak.com/meta-admits-use-of-pirated-book-dat...

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#12
As I've said several times, the corpus is key: LLMs thus far "read" most anything, but should instead have well-curated corpora. "Garbage In, Garbage Out!(GIGO)" is the saying.

While the Harry Potter series may be fun reading, it doesn't provide information about anything that isn't better covered elsewhere. Leave Harry Potter for a different "Harry Potter LLM".

Train scientific LLMs to the level of a good early 20th century English major and then use science texts and research papers for the remainder.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#13

As I've said several times, the corpus is key: LLMs thus far "read" most anything, but should instead have well-curated corpora. "Garbage In, Garbage Out!(GIGO)" is the saying. While the Harry Potter series may be fun reading, it doesn't provide information about anything that isn't better covered elsewhere. Leave Harry Potter for a different "Harry Potter LLM". Train scientific LLMs to the level of a good early 20th…

> While the Harry Potter series may be fun reading, it doesn't provide information about anything that isn't better covered elsewhere

It has copyright implications - if Claude can recollect 42% of a copyrighted product without attribution or royalties, how did Anthropic train it?

> Train scientific LLMs to the level of a good early 20th century English major and then use science texts and research papers for the remainder

Plenty of in-stealth companies approaching LLMs via this approach ;)

For those of us who studied the natural sciences and CS in the 2000s and early 2010s, there was a bit of a trend where certain PIs would simply translate German and Russian papers from the early-to-mid 20th century and attribute them to themselves in fields like CS (especially in what became ML).

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#15

As I've said several times, the corpus is key: LLMs thus far "read" most anything, but should instead have well-curated corpora. "Garbage In, Garbage Out!(GIGO)" is the saying. While the Harry Potter series may be fun reading, it doesn't provide information about anything that isn't better covered elsewhere. Leave Harry Potter for a different "Harry Potter LLM". Train scientific LLMs to the level of a good early 20th…

> While the Harry Potter series may be fun reading, it doesn't provide information about anything that isn't better covered elsewhere It has copyright implications - if Claude can recollect 42% of a copyrighted product without attribution or royalties, how did Anthropic train it? > Train scientific LLMs to the level of a good early 20th century English major and then use science texts and research papers for the rema…

Why are you talking about Claude and Anthropic?

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#16

As I've said several times, the corpus is key: LLMs thus far "read" most anything, but should instead have well-curated corpora. "Garbage In, Garbage Out!(GIGO)" is the saying. While the Harry Potter series may be fun reading, it doesn't provide information about anything that isn't better covered elsewhere. Leave Harry Potter for a different "Harry Potter LLM". Train scientific LLMs to the level of a good early 20th…

> While the Harry Potter series may be fun reading, it doesn't provide information about anything that isn't better covered elsewhere It has copyright implications - if Claude can recollect 42% of a copyrighted product without attribution or royalties, how did Anthropic train it? > Train scientific LLMs to the level of a good early 20th century English major and then use science texts and research papers for the rema…

So if I memorized Harry Potter the physical encoding which definitely exists in my brain is a copyright violation?

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#17

As I've said several times, the corpus is key: LLMs thus far "read" most anything, but should instead have well-curated corpora. "Garbage In, Garbage Out!(GIGO)" is the saying. While the Harry Potter series may be fun reading, it doesn't provide information about anything that isn't better covered elsewhere. Leave Harry Potter for a different "Harry Potter LLM". Train scientific LLMs to the level of a good early 20th…

That's got nothing to do with it. It's all about copyright. Can it reproduce its training data verbatim? If so, Meta is in hot water.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#18

Earlier quoted context omitted.

> While the Harry Potter series may be fun reading, it doesn't provide information about anything that isn't better covered elsewhere It has copyright implications - if Claude can recollect 42% of a copyrighted product without attribution or royalties, how did Anthropic train it? > Train scientific LLMs to the level of a good early 20th century English major and then use science texts and research papers for the rema…

So if I memorized Harry Potter the physical encoding which definitely exists in my brain is a copyright violation?

I think humans get a special exception in cases like this

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#19

Earlier quoted context omitted.

> While the Harry Potter series may be fun reading, it doesn't provide information about anything that isn't better covered elsewhere It has copyright implications - if Claude can recollect 42% of a copyrighted product without attribution or royalties, how did Anthropic train it? > Train scientific LLMs to the level of a good early 20th century English major and then use science texts and research papers for the rema…

So if I memorized Harry Potter the physical encoding which definitely exists in my brain is a copyright violation?

maybe if you re wrote it from memory.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#20

Earlier quoted context omitted.

> While the Harry Potter series may be fun reading, it doesn't provide information about anything that isn't better covered elsewhere It has copyright implications - if Claude can recollect 42% of a copyrighted product without attribution or royalties, how did Anthropic train it? > Train scientific LLMs to the level of a good early 20th century English major and then use science texts and research papers for the rema…

So if I memorized Harry Potter the physical encoding which definitely exists in my brain is a copyright violation?

You are not selling or distributing copies of your brain.
Post reply on HN