Live data from Hacker News

Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

understandingai.org

301–310 of 326 posts

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#301

Earlier quoted context omitted.

Everything is derivative. This boundary you are defending between originality and slop is extremely subjective at best. What harm is slop anyway? If originality is so objectively valuable, then why should its value be systemically enforced? At the intersection of capitalism and copyright, I see a serious problem. Collaboration is encapsulated by competition. Because simple derivative work is illegal, all collaboratio…

We know what the world looks like without copyright and that world has far fewer works created and very few artists who can do it full-time absent patronage or independent wealth. Banning the nonsense that is character copyright and shortening copyright back down to a reasonable length of time (say, 20 years) would still enable the creation of more culturally-relevant derivative works without pauperizing every artist…

How could we possibly know that? Copyright has existed since before the industrial revolution even started. What you described is not really that far from reality today: most artists are not really making a living. The words "starving artist" have not even begun to lose their meaning. Every artist I know has been failed by copyright. The value a copyright creates is not applied to the art: it's applied to the moat around the art. The only certain beneficiaries are the giant corporations that use their collected moats to drown out small competition, including artists.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#302
post #22

It's important to note the way it was measured: > the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time. So 50 tokens is not really…

Even if it is recalling it 50 tokens at a time, the half of the book is in some sense in there, right?

If I give you a vision algorithm that, given every other frame of a Harry Potter movie, can accurately predict the interstitials - would you say that half that Harry Potter movie is "in" it?

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#303
post #104

Earlier quoted context omitted.

Of course it is. It's just a form of compression. If I train an autoencoder on an image, and distribute the weights, that would obviously be the same as distributing the content. Just because the content is commingled with lots of other content doesn't make it disappear. Besides, where did the sections of text from the input works that show up in the output text come from? Divine inspiration? God whispering to the ma…

Indeed! It is a form of massive lossy compression. > Llama 3 70B was trained on 15 trillion tokens That's roughly a 200x "compression" ration; compared to 3-7x for tradtional lossless text compression like bzip and friends. LLM don't just compress, they generalize. If they could only recite Harry Potter perfectly but couldn’t write code or explain math, they wouldn’t be very useful.

But LLMs cant write code nor explain math, they only plagiarize existing code and plagiarize existing explanations of math.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#304

Earlier quoted context omitted.

We know what the world looks like without copyright and that world has far fewer works created and very few artists who can do it full-time absent patronage or independent wealth. Banning the nonsense that is character copyright and shortening copyright back down to a reasonable length of time (say, 20 years) would still enable the creation of more culturally-relevant derivative works without pauperizing every artist…

How could we possibly know that? Copyright has existed since before the industrial revolution even started. What you described is not really that far from reality today: most artists are not really making a living. The words "starving artist" have not even begun to lose their meaning. Every artist I know has been failed by copyright. The value a copyright creates is not applied to the art: it's applied to the moat ar…

The copyright laws that existed prior to the industrial revolution only existed only in a small number of countries. A large swath of the planet had no equivalent.

Even British Colonial America had no copyright, save a handful of exceptions, as the Statute of Anne did not apply to the colonies.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#305
post #302

Earlier quoted context omitted.

Even if it is recalling it 50 tokens at a time, the half of the book is in some sense in there, right?

If I give you a vision algorithm that, given every other frame of a Harry Potter movie , can accurately predict the interstitials - would you say that half that Harry Potter movie is "in" it?

Congratulations, you've just invented a video codec with motion estimation. The motion experts group wants their share on some bullshit royalties/patents though, better pay up because they are very litigious and won't go soft on you because you are not a big tech corporation :)

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#306

It will generate a correct next token 42% of the time when prompted with a 50 token quote. Not 42% of the book. It's a pretty big distinction.

That said, I bet that if you could lower the inference temperature such chances would improve by a lot.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#307
post #294

Earlier quoted context omitted.

Copyright fair use rules are tools designed to govern how humans use protected works in dervied works. AI is not human use, therefore the rules are only coincidentally correct for AI use where it even is.

If you take that approach to fair use, don't you open the door to the same argument for copyright itself? How do you distinguish between a tool and the director of a tool? I doubt people would say that a person is immune to copyright or fair use rules because it was the pen that wrote the document, not the person.

> don’t you open the door to the same argument for copyright itself?

Yes, it comes down to intentional control of output. Copyright applies when someone uses a pen to make a drawing because of the degree of control.

On the flip side there are copyright free photos where an animal picked up a camera etc, the same applies to a great deal of automatically generated data. The output of an LLM is likely in the public domain unless it’s a derivative work of something in the training set.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#308

Earlier quoted context omitted.

Training itself involves making infringing copies of protected works. Whether or not inference produces copyrighted material is almost beside the point.

It’s legal if it’s fair use, which is yet decided by court

[deleted]

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#309
harry potter is likely to be excerpted a million times all over the web (in legitimate fair use context). wouldn't it make more sense to try out other titles that are still under copyright, appear in the research datasets, but have little mention across the web and other typical source corpii?

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#310

Earlier quoted context omitted.

All of the major AI models these days use "clean" datasets stripped of copyrighted material. They also use data from the previous models, so I'm not sure how "clean" it really is

> All of the major AI models these days use "clean" datasets stripped of copyrighted material. Which of the major commercial models discloses its dataset? Or are you just trusting some unfalsifiable self-serving PR characterization?

It's from my personal experience in the industry
Post reply on HN