Live data from Hacker News

Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

understandingai.org

91–100 of 326 posts

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#91
post #74

Earlier quoted context omitted.

Did you read the article? This exact point is made and then analyzed. > Or maybe Meta added third-party sources—such as online Harry Potter fan forums, consumer book reviews, or student book reports—that included quotes from Harry Potter and other popular books. > “If it were citations and quotations, you'd expect it to concentrate around a few popular things that everyone quotes or talks about,” Lemley said. The fac…

The article fails to mention or understand the volume of content here. Every, literally every, part of these books is quoted and "talked about" (in the sense of used in unlicensed derivative works). And yes, I read the article before commenting. I don't appreciate the baseless insinuation to the contrary.

Agreed. It’s an obtuse quote by Lemley who can’t picture the enormous quantity of associations and crawled data, or at least wants to minimize the quantity. It’s hardly discussion-ending.

Accusations of not reading the article are fair when someone brings up a “related” anecdote that was in the article. It’s not fair when someone is just disagreeing.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#92

I really wish we could get rid of copyright. It's going to hold us back long term.

We cannot get ride of it without finding a way to pay the creators that generate copyrighted works. I’m personally more in favor of significantly reducing the length of the copy right. I think 20-30 years is an interesting range. Artist get roughly a career length of time to profit off their creations, but there is much less incentive for major corporations to buy and horde IP.

[deleted]

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#93
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

Indeed but since when is a blatantly derived work only using 50% of a copyrighted work without permission a paragon of copyright compliance? Music artists get in trouble for using more than a sample without permission — imagine if they just used 45% of a whole song instead… I’m amazed AI companies haven’t been sued to oblivion yet. This utter stupidity only continues because we named a collection of matrices “Artific…

> a blatantly derived work only using 50% of a copyrighted work without permission

What's the work here? If it's the output of the LLM, you have to feed in the entire book to make it output half a book so on an ethical level I'd say it's not an issue. If you start with a few sentences, you'll get back less than you put in.

If the work is the LLM itself, something you don't distribute is much less affected by copyright. Go ahead and play entire songs by other artists during your jam sessions.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#94

On the other hand, it’s surprising that Llama memorized so much of Harry Potter and the Sorcerer's Stone. It's sold 120 million copies over 30 years. I've gotta think literally every passage is quoted online somewhere else a bunch of times. You could probably stitch together the full book quote-by-quote.

If I collect HP quotes from the internet and then stitch them together into a book, can I legally sell access it?

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#95
post #25

Earlier quoted context omitted.

> the physical encoding which definitely exists in my brain is a copyright violation First of all, we don't really know how the brain works. I get that you're being a snarky physicalist, but there's plenty of substance dualists, panpsychsts, etc. out there. So, some might say, this is a reductive description of what happens in our brains. Second of all, yes, if you tried to publish Harry Potter ( even if it was from…

Right but the physical encoding already exists in my brain or how can I reproduce it in the first place? We may not know how the encoding works but we do know that an encoding exists because a decoding is possible. My question is… is that in itself a violation of copyright? If not then as long as LLMs don’t make a publication it shouldn’t be a copyright violation right? Because we don’t understand how it’s encoded in…

copyright is actually not as much about right to copy as it is about redistribution permissions.

if you trained an LLM on real copyrighted data, benchmarked it, wrote up a report, and then destroyed the weight, that's transformative use and legal in most places.

if you then put up that gguf on HuggingFace for anyone to download and enjoy, well... IANAL. But maybe that's a bit questionable, especially long term.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#96
post #83
post #75

Earlier quoted context omitted.

That may be relevant in the NYT vs OpenAI case, since NYT was supposedly able to reproduce entire articles in ChatGPT. Here Llama is predicting one sentence at a time when fed the previous one, with 50% accuracy, for 42% of the book. That can easily be written off as fair use.

I'm pretty sure books.google.com does the exact same with much better reliability... and the US courts found that to be fair use. (Agreeing with parent comment)

If there is a circuit split between it and NYT vs OAI, the Google Books ruling (in the famously tech-friendly ninth circuit) may also find itself under review.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#98
post #90
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

You really don't see the difference between Google indexing the content of third parties and directly hosting/distributing the content itself?

Where are they putting any blame on Google here?

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#99
post #90
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

You really don't see the difference between Google indexing the content of third parties and directly hosting/distributing the content itself?

Hosting model weights is not hosting / distributing the content.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#100
It’s well-known that John von Neumann had this ability too:

Herman Goldstine wrote "One of his remarkable abilities was his power of absolute recall. As far as I could tell, von Neumann was able on once reading a book or article to quote it back verbatim; moreover, he could do it years later without hesitation. He could also translate it at no diminution in speed from its original language into English. On one occasion I tested his ability by asking him to tell me how A Tale of Two Cities started. Whereupon, without any pause, he immediately began to recite the first chapter and continued until asked to stop after about ten or fifteen minutes."

Maybe it’s just an unavoidable side effect of extreme intelligence?

Post reply on HN