Live data from Hacker News

Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

understandingai.org

191–200 of 326 posts

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#191

Earlier quoted context omitted.

I think the argument is less about piracy and more that the model(s output) is a derivative work of Harry Potter, and the rights holder should be paid accordingly when it’s reproduced.

If the assertion in the parent comment is correct "nobody is using this as a substitute to buying the book" why should the rights holders get paid?

Facebook are using the contents of the book to make money.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#193
post #156

Earlier quoted context omitted.

> let's not pretend that an LLM that autocompletes a couple lines from harry potter with 50% accuracy is some massive new avenue to piracy No one is claiming this. The corporations developing LLMs are doing so by sampling media without their owners' permission and arguing this is protected by US fair use laws, which is incorrect - as the late AI researcher Suchir Balaji explained in this other article: https://suchir…

Yeah, that's literally the title of the article,and the premise of the first paragraph.

It's not literally the title of the article, nor the premise of its first paragraph, but since this was your interpretation I wonder if there is a misunderstanding around the term "piracy", which I believe is normally defined as the unauthorized reproduction of works, not a synonym for copyright infringement, which is a more broad concept.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#195
That's a clickbait title.

What they are actually saying: Given one correct quoted sentence, the model has 42% chance of predicting the next sentence correctly.

So, assuming you start with the first sentence and tell it to keep going, it has a 0.42^n odds of staying on track, where n is the n-th sentence.

It seems to me, that if they didn't keep correcting it over and over again with real quotes, it wouldn't even get to the end of the first page without descending into wild fanfiction territory, with errors accumulating and growing as the length of the text progressed.

EDIT: As the article states, for an entire 50 token excerpt to be correct the probability of each output has to be fairly high. So perhaps it would be more accurate to view it as 0.985^n where n is the n-th token. Still the same result long term. Unless every token is correct, it will stray further and further from the correct source.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#197
Imagine the literary possibilities when it can write 100%! Rowling's original work was an amusing, if rather derivative children's book. But Llama's version of the Philosophers stone will be something else entirely. Just think of the rather heavy-handed Cerberus reference in the original work. Instead of a rote reference to Greek mythology used as a simple trope, it will be filled with a subtext that only an LLM can produce.

Right now they're working on recreating the famous sequence with the troll in the dungeon. It might cost them another few billion in training, but the end results will speak for themselves.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#198
post #22

It's important to note the way it was measured: > the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time. So 50 tokens is not really…

The claim of the paper is not so much that the model is reproducing content illegally but that harry Potter has been used to train the model.

This does not appear to happen with other models they tested to the same degree

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#199
post #184
post #178

Earlier quoted context omitted.

I’ve yet to read an actual argument defending commercial LLM’s as fair use based on existing (edit:legal) criteria.

It seems like a pretty reasonable argument and easy enough to make. A human with a great memory could probably recreate some absurd % of Harry Potter after reading it, there are some very unusual minds out there. It is clear that if they read Harry Potter and being capable of reproducing it on demand as a party trick that would be fair use. So the LLM should also be fair use since it is using a mechanism similar enou…

> is clear that if they read Harry Potter and reproduce it on demand as a party trick that would be fair use.

Actually no that could be copyright infringement. Badly signing a recent pop song in public also qualifies as copyright infringement. Public performances count as copying here.

Post reply on HN