Live data from Hacker News

Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

understandingai.org

311–320 of 326 posts

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#311

Earlier quoted context omitted.

> All of the major AI models these days use "clean" datasets stripped of copyrighted material. Which of the major commercial models discloses its dataset? Or are you just trusting some unfalsifiable self-serving PR characterization?

It's from my personal experience in the industry

What are your thoughts on the origin of the LLaMA leak? It's interesting that the training data was torrented, and so was the leak. Perhaps we will never know? For the OSINT folks, not a lot to go on, or maybe a lot, depending?

https://en.wikipedia.org/wiki/Llama_(language_model)#Leak

https://archived.moe/g/thread/91848262#p91850335

https://github.com/meta-llama/llama/pull/73/files

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#312
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

It is actually much worse than piracy. I would much prefer a complete pirate copy of my creation to a half baked one.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#313
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

It is actually much worse than piracy. I would much prefer a complete pirate copy of my creation to a half baked one.

It's kind of a no-win situation for creators, as their work is bastardized and name divorced from all meaning it might have as a creator in relation to zombie necroposts sprung to life. One's own right to be identified as a creator is made meaningless in relation to such a creation that they didn't directly create. AI are apocalyptic plague locusts that convert coal to droll; AI are alienation demons driving human mothers and fathers against their own estranged reanimated lifeless intellectual prodigal child Frankensteins.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#314
post #156
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

> let's not pretend that an LLM that autocompletes a couple lines from harry potter with 50% accuracy is some massive new avenue to piracy No one is claiming this. The corporations developing LLMs are doing so by sampling media without their owners' permission and arguing this is protected by US fair use laws, which is incorrect - as the late AI researcher Suchir Balaji explained in this other article: https://suchir…

> The corporations developing LLMs are doing so by sampling media without their owners' permission and arguing this is protected by US fair use laws

The schools developing future labor are doing so by sampling media without their owners' permission ...

Harry Potter's been required reading for a while.

And while it may not be the most quoted, others are: any time we're supposed to understand a movie or TV character is "well read" they quote paragraphs from famous authors.

This is what libraries used to be for, and what the Internet was supposed to be for: fill our brains with what's been published and hopefully we remember some of it. Should libraries be off limits to savants with eidetic memory?

Why should learning from reading be off limits to the machine?

// Reproducing the reading material for distribution is illegal for both man and machine.

However, perhaps you're using a different definition of "use" in fair use, than the traditional "quote it".

- the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and

- the effect of the use upon the potential market for or value of the copyrighted work.

In general these are not thought to mean you can't use it as in learn from it, they are thought to mean you can't reproduce chunks, perform chunks, etc.

I imagined it's well established that "learn" is not "use".

So then where I find myself uncertain is whether learning, then responding about it, is learning + (hand waving) artificial intelligence, or whether it's just source (context) compression with prompted continuation to mine the context, and what density of words from the source in the continuation starts to be "use".

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#315

Earlier quoted context omitted.

It's from my personal experience in the industry

What are your thoughts on the origin of the LLaMA leak? It's interesting that the training data was torrented, and so was the leak. Perhaps we will never know? For the OSINT folks, not a lot to go on, or maybe a lot, depending? https://en.wikipedia.org/wiki/Llama_(language_model)#Leak https://archived.moe/g/thread/91848262#p91850335 https://github.com/meta-llama/llama/pull/73/files

I don't really know much about that, sorry

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#316

Earlier quoted context omitted.

What are your thoughts on the origin of the LLaMA leak? It's interesting that the training data was torrented, and so was the leak. Perhaps we will never know? For the OSINT folks, not a lot to go on, or maybe a lot, depending? https://en.wikipedia.org/wiki/Llama_(language_model)#Leak https://archived.moe/g/thread/91848262#p91850335 https://github.com/meta-llama/llama/pull/73/files

I don't really know much about that, sorry

I didn’t ask for info, I asked for your views. I gave you all the info anyone has publicly, so you have enough to comment.

I suspect that it was a limited hangout self-own by Meta to claim that they aren’t responsible, and then they are doing research on a leaked LLM that they developed, but then was leaked, so they can claim that the subsequent research is not tainted by the fruit of the poisonous tree legal doctrine. Or, their torrent client or other software on the same machine had 0-days and they got hacked by someone on the Books3 swarm or knowledgeable of what IPs were connecting to it.

I appreciate your posts and I am replying to you to humbly ask you to post more. :P

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#317

Earlier quoted context omitted.

I don't like many things about this post, its a bit snobbish and uses esoteric language in order to sound more intricate than it really is. >Capitalism is not simply incapable of reflection. In fact, it's structured to ignore it. It has no native interest in what emerges from its aggregated behaviors unless those emergent properties threaten the throughput of capital itself. It isn't designed to ask, "What kind of so…

Appreciate the engagement. But your reply mostly recenters a pro-capitalist narrative by redefining the products of "emergence" as inherently good. My argument isn't about stacking pros and cons and calculating the combined sum. It’s about a structural blind spot: capitalism systematically collapses higher-order questions about what kind of world were building into first-order value propositions like "growth," "utili…

There is a problem with your argument here:

>collapses higher-order questions about what kind of world were building into first-order value

But then

>you'd interrogate not what the system produces, but what it refuses to ask, and who gets to decide. That silence is precisely my point.

The reply your interlocutor provided is aligned with your incoherence. In a first move you point out that capitalism flattens everything into first-order land, and yet in a second move you tell us there are things it can't talk about. I guess your silence is precisely what articulates these two aspects of your discourse.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#318

Earlier quoted context omitted.

I don't really know much about that, sorry

I didn’t ask for info, I asked for your views. I gave you all the info anyone has publicly, so you have enough to comment. I suspect that it was a limited hangout self-own by Meta to claim that they aren’t responsible, and then they are doing research on a leaked LLM that they developed, but then was leaked, so they can claim that the subsequent research is not tainted by the fruit of the poisonous tree legal doctrin…

I'm not really sure what you are insinuating? You think Meta leaked LAMA so they could claim, legally that they are in the clear for copyright violation? Sorry, I just don't really get what you want me to opine about.

If that is what you are asking, I don't think that's what happened. It's far more likely that it was just leaked or grabbed by a hacker

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#319

Earlier quoted context omitted.

If you distribute a zip file of the book, are you violating copyright, or is it the person who unzips it?

If you walk through the N-gram database with a copy of Harry Potter in hand and observe that for N=7, you can find any piece of it in the database with above-average frequency, does that mean N-gram database is violating copyright?

Not unless you can reproduce large portions of Harry Potter verbatim from the database. If the 7-grams are taken only from Harry Potter, that is very likely.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#320

Earlier quoted context omitted.

I keep waiting for the day when people realise that IP law has been used and abused and thanks to Disney extended out for many, many lifetimes and all manner of dirty tricks/hacks to keep the late stage capitalism profit engine going. I 100% agree that if an LLM can entirely reproduce a book then that is copyright infringement, overfitting and generally a bad model. I also believe that in this case, HP (and other pop…

Abuse of IP does not mean the law is not relevant. Don’t you find it funny that corporations with market caps the size of small countries first sued people for singing Happy Birthday in public, and now pretend that IP is suddenly not really a thing? Do you really want to defend their interests?

I most certainly would like to see tax havens abolished worldwide. And I would most certainly like to see these corporations and their executive pay the taxes that they should be paying.

But I've come to realise that the apathetic general public only care about being racist, sexist and homophobic to each other whilst the ruling class laugh all the way to the bank. And the few intelligent people who understand what's going on don't raise their voices so long as their high tech salary is protected, so long as their taxes aren't raised, so long as they can own a holiday home or two when others struggle for their first home.

If you actually look into the numbers, what tax companies like Apple, Google, Amazon etc actually pay it's just...yeah.

Post reply on HN