Live data from Hacker News

Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

understandingai.org

271–280 of 326 posts

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#271

Earlier quoted context omitted.

If you distribute a zip file of the book, are you violating copyright, or is it the person who unzips it?

If you walk through the N-gram database with a copy of Harry Potter in hand and observe that for N=7, you can find any piece of it in the database with above-average frequency, does that mean N-gram database is violating copyright?

If the database is sharing those pieces, it might be yes.

Copyright takes into account the use for such the copying is done. Commercial use will almost always be treated as not fair use, with limited exceptions.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#272
post #253

Earlier quoted context omitted.

I keep waiting for the day when software stops being compared to a human person (a being with agency, free will, consciousness, and human rights of its own) for the purposes of justifying IP law circumvention. Yes, there is no problem when a person reads some book and recalls pieces[0] of it in a suitable context. How would that in any way address when certain people create and distribute commercial software, providi…

> I keep waiting for the day when software stops being compared to a human person (a being with agency, free will, consciousness, and human rights of its own) for the purposes of justifying IP law circumvention. I mean, "agency" is a goal of some AI; "free will" is incoherent*; the word "consciousness" has about 40 different definitions, some of which are so broad they include thermostats and others so narrow that it…

> and "human rights" are a purely legal concept.

Yep, and if you claim that a thing can reproduce IP like a human then you should explain why you are also not holding its operators to the same legal standard (try to use a human in the same way and it will be considered torture and slavery).

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#273

Earlier quoted context omitted.

I keep waiting for the day when software stops being compared to a human person (a being with agency, free will, consciousness, and human rights of its own) for the purposes of justifying IP law circumvention. Yes, there is no problem when a person reads some book and recalls pieces[0] of it in a suitable context. How would that in any way address when certain people create and distribute commercial software, providi…

I keep waiting for the day when people realise that IP law has been used and abused and thanks to Disney extended out for many, many lifetimes and all manner of dirty tricks/hacks to keep the late stage capitalism profit engine going. I 100% agree that if an LLM can entirely reproduce a book then that is copyright infringement, overfitting and generally a bad model. I also believe that in this case, HP (and other pop…

Abuse of IP does not mean the law is not relevant.

Don’t you find it funny that corporations with market caps the size of small countries first sued people for singing Happy Birthday in public, and now pretend that IP is suddenly not really a thing? Do you really want to defend their interests?

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#274
post #75

Earlier quoted context omitted.

That may be relevant in the NYT vs OpenAI case, since NYT was supposedly able to reproduce entire articles in ChatGPT. Here Llama is predicting one sentence at a time when fed the previous one, with 50% accuracy, for 42% of the book. That can easily be written off as fair use.

> Here Llama is predicting one sentence at a time when fed the previous one, with 50% accuracy, for 42% of the book. That can easily be written off as fair use. Is that fair use, or is that compression of the verbatim source?

It doesn't let you recover the text without knowing it in advance, so no.

You can't in particular iterate it sentence by sentence; you're unlikely to go past sentence 2 this way before it starts giving you back it's own ideas.

The whole thing is a sleigh of hand, basically. There's 42% of the book there, in tiny pieces, which you can only identify if you know what you're looking for. The model itself does not.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#275
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

[dead]

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#276
post #215

Earlier quoted context omitted.

Based upon legal decisions in the past there is a clear argument that the distinction for fair use is whether a work is substantially different to another. You are allowed to write a book containg information you learned about from another book. There is threshold in academia regarding plagiarism that stands apart from the legal standing. The measure that was used in Gyles v Wilcox was if the new work could substitut…

Training itself involves making infringing copies of protected works. Whether or not inference produces copyrighted material is almost beside the point.

No it doesn’t? You can buy a digital copy of Harry Potter and use it for training. No infringement needed.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#277
post #276

Earlier quoted context omitted.

Training itself involves making infringing copies of protected works. Whether or not inference produces copyrighted material is almost beside the point.

No it doesn’t? You can buy a digital copy of Harry Potter and use it for training. No infringement needed.

[deleted]

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#278
post #276

Earlier quoted context omitted.

Training itself involves making infringing copies of protected works. Whether or not inference produces copyrighted material is almost beside the point.

No it doesn’t? You can buy a digital copy of Harry Potter and use it for training. No infringement needed.

Only as long as it's not copied again during training. You can't make copies of your purchased digital copy for any reason other than archival.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#279
post #178

Earlier quoted context omitted.

It's not clear that it's incorrect.

I’ve yet to read an actual argument defending commercial LLM’s as fair use based on existing (edit:legal) criteria.

"It's just doing what a human would do!" -Internet AI Expert

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#280

Earlier quoted context omitted.

If you walk through the N-gram database with a copy of Harry Potter in hand and observe that for N=7, you can find any piece of it in the database with above-average frequency, does that mean N-gram database is violating copyright?

If the database is sharing those pieces, it might be yes. Copyright takes into account the use for such the copying is done. Commercial use will almost always be treated as not fair use, with limited exceptions.

I'd say no, because you can't reasonably access and order those pieces without already having the work at your side to use as a reference.
Post reply on HN