Live data from Hacker News

Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

understandingai.org

141–150 of 326 posts

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#141
post #137
post #133

Earlier quoted context omitted.

> Have you ever repeated a line from your favorite movie or TV show? Memorized a poem? Guess the rights holders better sue you for stealing their content by encoding it in your wetware neural network. I see this absolute non-argument regurgitated ad infinitum in every single discussion on this topic, and at this point I can't help but wonder: doesn't it say more about the person who says it than anything else? Do you…

Who said anything about freedom of speech? Nobody is claiming the LLM has free speech rights, which don't even apply to infringing copyright anyway. Freedom of speech doesn't give me the right to make copies of copyrighted works. The question is whether the model weights constitute of copy of the work. I contend that they do not, or they did, than so do the analogous weights (reinforced neural pathways) in your brain…

> Freedom of speech doesn't give me the right to make copies of copyrighted works.

No, but it gives you the right to quote a line from a movie or TV show without being charged with copyright infringement. You argued that an LLM deserves that same right, even if you didn't realize it.

> than so do the analogous weights (reinforced neural pathways) in your brain

Did your brain consume millions of copyrighted books in order to develop into what it is today? Would your brain be unable to exist in its current form if it had not consumed those millions of books?

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#142
post #134

Would it be possible that other people posted content of Harry Potter book online and the model developer scrape that information? Would the model developer be at fault in this scenario?

I think this is good question. At least for LLMs in general. However we know that Meta used pirated torrents.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#143
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

People aren't buying Harry Potter action figures as a subtitute for buying the book either, but copyright protects creators from other people swooping in and using their work in other mediums. There is obviously a huge market demand for high quality data for training LLMs, Meta just spent 15 billion on a data labeling company. Companies training LLMs on copyrighted material without permission are doing that as a substitue for obtaining a license from the creator for doing so in the same way that a pirate downloading a torrent is a substitue for getting an ebook license.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#144
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

So? Am I allowed to also ignore certain laws if I can prove others have also ignored them?

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#145
post #109
post #104

Earlier quoted context omitted.

Of course it is. It's just a form of compression. If I train an autoencoder on an image, and distribute the weights, that would obviously be the same as distributing the content. Just because the content is commingled with lots of other content doesn't make it disappear. Besides, where did the sections of text from the input works that show up in the output text come from? Divine inspiration? God whispering to the ma…

[flagged]

I have, but I never tried to make any money off of it either

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#146
post #141
post #137

Earlier quoted context omitted.

Who said anything about freedom of speech? Nobody is claiming the LLM has free speech rights, which don't even apply to infringing copyright anyway. Freedom of speech doesn't give me the right to make copies of copyrighted works. The question is whether the model weights constitute of copy of the work. I contend that they do not, or they did, than so do the analogous weights (reinforced neural pathways) in your brain…

> Freedom of speech doesn't give me the right to make copies of copyrighted works. No, but it gives you the right to quote a line from a movie or TV show without being charged with copyright infringement. You argued that an LLM deserves that same right, even if you didn't realize it. > than so do the analogous weights (reinforced neural pathways) in your brain Did your brain consume millions of copyrighted books in o…

Millions? No, but my brain certainly consumed thousands of books, movies, TV shows, pieces of music, artworks, and other copyrighted material. Where is the cutoff? Can I only consume 999,999 copyrighted works before I'm not longer allowed to remember something without infringing copyright? My brain definitely would not exist in its current form without consuming that material. It would exist in some form, but it would without a doubt be different than it is having consumed the material.

An LLM is not a person and does not deserve any rights. People have rights, including the right to use tools like LLMs without having to grease the palm of every grubby rights holder (or their great-great-grandchild) just because it turns out their work was so trite and predictable it could be reproduced by simply guessing the next most likely token.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#147
post #10

It will generate a correct next token 42% of the time when prompted with a 50 token quote. Not 42% of the book. It's a pretty big distinction.

This means that if we start with 50% of the book then there is 42% chance that we can recreate the remaining 50%. What is the distinction between understanding and memorization? What is the chance that understanding results in memorization (may be in case of humans)?

It stores how often characters will come next based on how often they happen in copyright material. It can reproduce parts because those values are a fingerprint.

It should break copyright laws as written now but too much money involved.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#148
post #75

Earlier quoted context omitted.

I think the argument is less about piracy and more that the model(s output) is a derivative work of Harry Potter, and the rights holder should be paid accordingly when it’s reproduced.

That may be relevant in the NYT vs OpenAI case, since NYT was supposedly able to reproduce entire articles in ChatGPT. Here Llama is predicting one sentence at a time when fed the previous one, with 50% accuracy, for 42% of the book. That can easily be written off as fair use.

That can easily be written off as fair use.

No, it really couldn't. In fact, it's very persuasive evidence that Llama is straight up violating copyright.

It would be one thing to be able to "predict" a paragraph or two. It's another thing entirely to be able to predict 42% of a book that is several hundred pages long.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#149
It's not fair use just because you guys want it be fair use.

While limited quoting can (and usually is) considered fair use, quoting significant portions of a book (much less 42% of it) has never been fair use, in the U.S., Europe, or any other nation.

Yes, information wants to be free, yada yada. That means facts. Whether creative works are free is up to their creators.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#150
post #22

It's important to note the way it was measured: > the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time. So 50 tokens is not really…

Even if it is recalling it 50 tokens at a time, the half of the book is in some sense in there, right?

I don’t think this paper proves that, and I don’t think it is in a traditional sense.

It can produce the next sentence or two, but I suspect it can’t reproduce anything like the whole text. If you were to recursively ask for the next 50 tokens, the first time it’s wrong the output would probably cease matching because you fed it not-Harry-Potter.

It seems like chopping Harry Potter up into 2 sentences at a time on post it’s and tossing those in the air. It does contain Harry Potter, in a way, but without the structure is it actually Harry Potter?

Post reply on HN