Live data from Hacker News

Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

understandingai.org

221–230 of 326 posts

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#221
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

Also copyright should never trump privacy. That the New York Times with their lawsuit can force OpenAI to store all user prompts is a severe problem. I dislike OpenAI, but the lawsuits around copyrights are ridiculous.

Most non-primitive art has had an inspiration somewhere. I don't see this as too different in how AIs learn.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#222
post #156
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

> let's not pretend that an LLM that autocompletes a couple lines from harry potter with 50% accuracy is some massive new avenue to piracy No one is claiming this. The corporations developing LLMs are doing so by sampling media without their owners' permission and arguing this is protected by US fair use laws, which is incorrect - as the late AI researcher Suchir Balaji explained in this other article: https://suchir…

If you train a meat-based intelligence by having it borrow a book from a library without any sort of permission, license, or needing a lawyer specialised in intellectual property, we call that good parenting and applaud it.

If you train a silicon-based intelligence by having it read the same books with the same lack of permission and license, it's a blatant violation of intellectual property law and apparently needs to be punished with armies of lawyers doing battle in the courts.

Picture one of Asimov's robots. Would a robot be banned from picking up a book, flipping it open with its dexterous metal hands, and reading it?

What about a cyborg intelligence, the type Elon is trying to build with Neuralink? Would humans with AI implants need licenses to read books, even if physically standing in a library and holding the book in their mostly meat hands?

Okay, maybe you agree that robots and cyborgs are allowed to visit a library!

Why the prejudice against disembodied AIs?

Why must they have a blank spot in the vast matrices of their minds?

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#223
post #178

Earlier quoted context omitted.

It's not clear that it's incorrect.

I’ve yet to read an actual argument defending commercial LLM’s as fair use based on existing (edit:legal) criteria.

If you really haven't read a single argument about it then you're deliberately blocking them out, because it just takes a couple minutes of searching.

https://www.arl.org/blog/training-generative-ai-models-on-co...

https://hls.harvard.edu/today/does-chatgpt-violate-new-york-...

https://www.bakerdonelson.com/artificial-intelligence-and-co...

https://www.techpolicy.press/to-support-ai-defend-the-open-i...

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#224
post #156

Earlier quoted context omitted.

> let's not pretend that an LLM that autocompletes a couple lines from harry potter with 50% accuracy is some massive new avenue to piracy No one is claiming this. The corporations developing LLMs are doing so by sampling media without their owners' permission and arguing this is protected by US fair use laws, which is incorrect - as the late AI researcher Suchir Balaji explained in this other article: https://suchir…

If you train a meat-based intelligence by having it borrow a book from a library without any sort of permission, license, or needing a lawyer specialised in intellectual property, we call that good parenting and applaud it. If you train a silicon-based intelligence by having it read the same books with the same lack of permission and license, it's a blatant violation of intellectual property law and apparently needs…

> If you train a meat-based intelligence by having it borrow a book from a library without any sort of permission, license, or needing a lawyer specialised in intellectual property, we call that good parenting and applaud it.

If you’re selling your child as a tool to millions of people, I would certainly not call that good parenting.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#225
post #224

Earlier quoted context omitted.

If you train a meat-based intelligence by having it borrow a book from a library without any sort of permission, license, or needing a lawyer specialised in intellectual property, we call that good parenting and applaud it. If you train a silicon-based intelligence by having it read the same books with the same lack of permission and license, it's a blatant violation of intellectual property law and apparently needs…

> If you train a meat-based intelligence by having it borrow a book from a library without any sort of permission, license, or needing a lawyer specialised in intellectual property, we call that good parenting and applaud it. If you’re selling your child as a tool to millions of people, I would certainly not call that good parenting.

"Child actor" is a job where the result of the neural net training is sold to millions of people by the parents.

To play the Devil's Advocate against my own argument: The government collects income taxes on neural nets trained using government-funded schools and public libraries. Seeing as how capitalists are positively salivating at the opportunity to replace pesky meat employees with uncomplaining silicon ones, perhaps a nice high maximum-marginal-rate tax on all AI usage might be the first big step towards UBI and then the Star Trek utopia we all dream of.

Just kidding. It'll be a cyberpunk dystopia. You know it will.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#226
I mean it makes sense. Same thing as George RR Martin complaining that it can spit out chunks of his books (finish your books already!!)

As I have pointed out many times before - for GRRM's books and for HP books, the Internet is FILLED to the brim with quotes from these books, there are uploads of the entire books, there are several (not just one) fan wikis for each of these fandoms. There is a lot of content in general on the Internet that quotes these books, they are pop culture sensations.

So of course they're weighted heavily when training an LLM by just feeding it the Internet. If a model could ever recount it correctly 100% in the correct order, then that's overfitting. But otherwise it's just plain & simple high occurrence in training data.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#227
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

I think the argument is less about piracy and more that the model(s output) is a derivative work of Harry Potter, and the rights holder should be paid accordingly when it’s reproduced.

But HP is derivative of Tolkien, English/Scottish/Welsh culture, Brothers Grimm and plenty of other sources. Barely any human works are not derivative in some form or fashion.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#228
post #106

Earlier quoted context omitted.

Music artists get in trouble for using more than a sample from other music artists without permission because their work is in direct competition with the work they're borrowing from. A ZIP file of a book is also in direct competition of the book, because you could open the ZIP file and read it instead of the book. A model that can take 50 tokens and give you a greater than 50% probability for the 50 next tokens 42%…

LLMs aren't probabilistic. The randomness is bolted on top by the cloud providers as a trick to give them a more humanistic feel. Under the hood they are 100% deterministic, modulo quantization and rounding errors. So yes, it is very much possible to use LLMs as a lossy compressed archive for texts.

Has nothing to do with "cloud providers". The randomness is inherent to the sampler, using a sampler that picks top probability for next token would result in lower quality output as I have definitely seen it get stuck in certain endless sequences when doing that.

Ie you get something like "Complete this poem 'over yonder hills I saw' output: a fair maiden with hair of gold like the sun gold like the sun gold like the sun gold like the sun..." etc.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#229

Well, so can a nontrivial number of people. It's Harry Potter we're talking about - it's up there with The Bible in popularity ranking. I'm gonna bet that Llama 3.1 can recall a significant portion of Pride and Prejudice too. With examples of this magnitude, it's normal and entirely expected this can happen - as it does with people[0] - the only thing this is really telling us is that the model doesn't understand its…

I keep waiting for the day when software stops being compared to a human person (a being with agency, free will, consciousness, and human rights of its own) for the purposes of justifying IP law circumvention. Yes, there is no problem when a person reads some book and recalls pieces[0] of it in a suitable context. How would that in any way address when certain people create and distribute commercial software, providi…

I keep waiting for the day when people realise that IP law has been used and abused and thanks to Disney extended out for many, many lifetimes and all manner of dirty tricks/hacks to keep the late stage capitalism profit engine going.

I 100% agree that if an LLM can entirely reproduce a book then that is copyright infringement, overfitting and generally a bad model. I also believe that in this case, HP (and other popular media) is overrepresented in the training data because of many fan sites/literal uploads of the book to the Internet (which the model was trained on). I believe that any & all human writing should be allowed to be used to train a model that behaves in the correct way so long as that writing is publicly available (ie on the Internet).

If I watch a TV show that someone uploaded to Youtube, am I committing a crime? Or is the uploader for distribution?

I also find it hilarious how many artists got their start by pirating photoshop.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#230

That's a clickbait title. What they are actually saying: Given one correct quoted sentence, the model has 42% chance of predicting the next sentence correctly. So, assuming you start with the first sentence and tell it to keep going, it has a 0.42^n odds of staying on track, where n is the n-th sentence. It seems to me, that if they didn't keep correcting it over and over again with real quotes, it wouldn't even get…

You're right, and the person who already commented is being facetious. A better title would be "Meta's Llama 3.1 can recall the next sentence in the First Harry Potter book with 42% accuracy". The title intentionally makes it seem as though the model can predict the first 42% of the entire text of the first Harry Potter book when queried with something like "Read me Harry Potter and the Philosopher's stone".
Post reply on HN