Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
understandingai.org
Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
1–10 of 326 posts
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#2If you've seen as many magnet links as I have, with your subconscious similarly primed with the foreknowledge of Meta having used torrents to download/leech (and possibly upload/seed) the dataset(s) to train their LLMs, you might scroll down to see the first picture in this article from the source paper, and find uncanny the resemblance of the chart depicted to a common visual representation of torrent block download status.
Can't unsee it. For comparison (note the circled part):
https://superuser.com/questions/366212/what-do-all-these-dow...
Previously, related:
Extracting memorized pieces of books from open-weight language models - https://news.ycombinator.com/item?id=44108926 - May 2025
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#3https://archive.is/OSQt6 If you've seen as many magnet links as I have, with your subconscious similarly primed with the foreknowledge of Meta having used torrents to download/leech (and possibly upload/seed) the dataset(s) to train their LLMs, you might scroll down to see the first picture in this article from the source paper, and find uncanny the resemblance of the chart depicted to a common visual representation…
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#4Not 42% of the book.
It's a pretty big distinction.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#5It will generate a correct next token 42% of the time when prompted with a 50 token quote. Not 42% of the book. It's a pretty big distinction.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#6It will generate a correct next token 42% of the time when prompted with a 50 token quote. Not 42% of the book. It's a pretty big distinction.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#7Guess the next word: Not all heros wear _____
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#8Could it be plausible that an LLM had ingested parts of the book via scrapping web pages like this and not the full copyrighted book and get results similar to those of the linked study?
[1] https://www.goodreads.com/work/quotes/4640799-harry-potter-a...
[2] ~30 portions x 68 pages
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#9Given the method and how the english language works, isn't that the expected outcome for any text that isnt highly technical? Guess the next word: Not all heros wear _____
https://en.wikipedia.org/wiki/Artificial_intelligence_and_co...
> See for example OpenAI's comment in the year of GPT-2's release: OpenAI (2019). Comment Regarding Request for Comments on Intellectual Property Protection for Artificial Intelligence Innovation (PDF) (Report). United States Patent and Trademark Office. p. 9. PTO–C–2019–0038. “Well-constructed AI systems generally do not regenerate, in any nontrivial portion, unaltered data from any particular work in their training corpus”
https://copyrightalliance.org/kadrey-v-meta-hearing/
> During the hearing, Judge Chhabria said that he would not take into account AI licensing markets when considering market harm under the fourth factor, indicating that AI licensing is too “circular.” What he meant is that if AI training qualifies as fair use, then there is no need to license and therefore no harmful market effect.
I know this is arguing against the point that this copyright lobbyist is making, but I hope so much that this is the case. The “if you sample, you must license” precedent was bad, and it was an unfair taking from the commons by copyright holders, imo.
The paper this post is referencing is freely available:
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#10It will generate a correct next token 42% of the time when prompted with a 50 token quote. Not 42% of the book. It's a pretty big distinction.
What is the distinction between understanding and memorization? What is the chance that understanding results in memorization (may be in case of humans)?