Earlier quoted context omitted.
I think the argument is less about piracy and more that the model(s output) is a derivative work of Harry Potter, and the rights holder should be paid accordingly when it’s reproduced.
If the assertion in the parent comment is correct "nobody is using this as a substitute to buying the book" why should the rights holders get paid?
Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
191–200 of 326 posts
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#192edit: never mind, I’ll just ask ChatGPT
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#193Earlier quoted context omitted.
> let's not pretend that an LLM that autocompletes a couple lines from harry potter with 50% accuracy is some massive new avenue to piracy No one is claiming this. The corporations developing LLMs are doing so by sampling media without their owners' permission and arguing this is protected by US fair use laws, which is incorrect - as the late AI researcher Suchir Balaji explained in this other article: https://suchir…
Yeah, that's literally the title of the article,and the premise of the first paragraph.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#194Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#195What they are actually saying: Given one correct quoted sentence, the model has 42% chance of predicting the next sentence correctly.
So, assuming you start with the first sentence and tell it to keep going, it has a 0.42^n odds of staying on track, where n is the n-th sentence.
It seems to me, that if they didn't keep correcting it over and over again with real quotes, it wouldn't even get to the end of the first page without descending into wild fanfiction territory, with errors accumulating and growing as the length of the text progressed.
EDIT: As the article states, for an entire 50 token excerpt to be correct the probability of each output has to be fairly high. So perhaps it would be more accurate to view it as 0.985^n where n is the n-th token. Still the same result long term. Unless every token is correct, it will stray further and further from the correct source.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#196It would be nice to know that at least our literature might survive the technological singularity.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#197Right now they're working on recreating the famous sequence with the troll in the dungeon. It might cost them another few billion in training, but the end results will speak for themselves.
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#198It's important to note the way it was measured: > the paper estimates that Llama 3.1 70B has memorized 42 percent of the first Harry Potter book well enough to reproduce 50-token excerpts at least half the time As I understand it, it means if you prompt it with some actual context from a specific subset that is 42% of the book, it completes it with 50 tokens from the book, 50% of the time. So 50 tokens is not really…
This does not appear to happen with other models they tested to the same degree
Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
#199Earlier quoted context omitted.
I’ve yet to read an actual argument defending commercial LLM’s as fair use based on existing (edit:legal) criteria.
It seems like a pretty reasonable argument and easy enough to make. A human with a great memory could probably recreate some absurd % of Harry Potter after reading it, there are some very unusual minds out there. It is clear that if they read Harry Potter and being capable of reproducing it on demand as a party trick that would be fair use. So the LLM should also be fair use since it is using a mechanism similar enou…
Actually no that could be copyright infringement. Badly signing a recent pop song in public also qualifies as copyright infringement. Public performances count as copying here.