Live data from Hacker News

Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

understandingai.org

291–300 of 326 posts

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#291

Earlier quoted context omitted.

Even if it is recalling it 50 tokens at a time, the half of the book is in some sense in there, right?

Almost the entire book is in there. From the paper, if you give it a 100 token prompt, it will produce the next 50 tokens with more than 1% probability so that the produced tokens cover 91% of the book. And as the title says, it also produces next 50 tokens with more than 50% probability, so produced tokens cover 42% of the book. Bet it gets close to 100% as you reduce the probability. Also they went through the book…

I think I agree with this take. The book is in there in some sense, whether or not it is a copyright violation is debatable.

Honestly, I get why these debates happen—it is practical to establish whether or not this emerging tech is illegal under current law. But it’s also like… well, obviously current law wasn’t written with this sort of application in mind.

Whether or not we think LLMs are basically good or bad, they are clearly quite impactful. It would be a nice time to have a functional legislature to address this directly.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#292
post #290

Earlier quoted context omitted.

> The point is that "human rights" are not part of the natural order, they only exist as laws. I have seen this argument used before; when something suits an argument, it’s “nature”, when something doesn’t then it isn’t. I think it’s a fallacy. Humans are part of natural order. Our laws and how they evolve are part of our nature, and by extension part of nature. I don’t believe in silencing such discussions as irrele…

By that tautology, so are rocks. Rocks don't get natural rights. Also, you may note that "human rights" is a recent invention and not actually enforced worldwide even today.

> you may note that "human rights" is a recent invention and not actually enforced worldwide even today.

Consider that countries known for stronger interpretation of human rights and freedoms, including intellectual property rights, are also the countries at the forefront of innovation, including technical innovation that laid the foundation for LLMs in the first place. I think that is not a coincidence, and we should keep it in mind when there is a push to be dismissive of these concepts (which predominantly serves the interests of commercial LLM operators and their supply chain).

I’m sure you would not argue from a point where this recent interpretation of human rights is bad or incorrect, but if you would then perhaps there’s not much of a constructive discussion to be had. I would still oppose the use of the natural vs. unnatural distinction as the basis of that argument, though.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#293
post #17

As I've said several times, the corpus is key: LLMs thus far "read" most anything, but should instead have well-curated corpora. "Garbage In, Garbage Out!(GIGO)" is the saying. While the Harry Potter series may be fun reading, it doesn't provide information about anything that isn't better covered elsewhere. Leave Harry Potter for a different "Harry Potter LLM". Train scientific LLMs to the level of a good early 20th…

That's got nothing to do with it. It's all about copyright. Can it reproduce its training data verbatim? If so, Meta is in hot water.

But if it's corpora do NOT include the Harry Potter books then Meta is NOT in hot water,! So take the Harry Potter books out of the corpora. What is lost? Nothing IMO useful other than the ability to discuss Harry Potter books. BFD.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#294
post #215

Earlier quoted context omitted.

Based upon legal decisions in the past there is a clear argument that the distinction for fair use is whether a work is substantially different to another. You are allowed to write a book containg information you learned about from another book. There is threshold in academia regarding plagiarism that stands apart from the legal standing. The measure that was used in Gyles v Wilcox was if the new work could substitut…

Copyright fair use rules are tools designed to govern how humans use protected works in dervied works. AI is not human use, therefore the rules are only coincidentally correct for AI use where it even is.

If you take that approach to fair use, don't you open the door to the same argument for copyright itself?

How do you distinguish between a tool and the director of a tool? I doubt people would say that a person is immune to copyright or fair use rules because it was the pen that wrote the document, not the person.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#295
post #290

Earlier quoted context omitted.

By that tautology, so are rocks. Rocks don't get natural rights. Also, you may note that "human rights" is a recent invention and not actually enforced worldwide even today.

> you may note that "human rights" is a recent invention and not actually enforced worldwide even today. Consider that countries known for stronger interpretation of human rights and freedoms, including intellectual property rights, are also the countries at the forefront of innovation, including technical innovation that laid the foundation for LLMs in the first place. I think that is not a coincidence, and we shoul…

> Consider that countries known for stronger interpretation of human rights and freedoms, including intellectual property rights, are also the countries at the forefront of innovation, including technical innovation that laid the foundation for LLMs in the first place. I think that is not a coincidence, and we should keep it in mind when there is a push to be dismissive of these concepts (which predominantly serves the interests of commercial LLM operators and their supply chain).

Several fallacies there.

First because China and the USA are opposite ends of the spectrum on many of the ways "freedom" is measured, and yet China is doing pretty well on the innovation front, including with AI. And that China is beating Europe, even though various Scandinavian nations rank higher on such freedoms than does the USA.

Second, cum hoc ergo propter hoc: Correlation does not imply causation. For example in this case, a reason why one of the big IP groups in the USA (Hollywood) got big, was because being in California enabled them to avoid the IP rights of the Motion Picture Patents Company that dominated cinema in the East Coast. I would even suggest that it is the disregarding of IP rights that enables much of the web, not only how and why China is doing well, but also Google (which has had legal fights over the interaction between copyright and search results), social media, and cultural elements such as memes and reaction gifs.

Third: the point of copyright is to encourage new works, because this makes money which can be taxed. All this becomes somewhat irrelevant when AI can also create new works.

If you want to set a bar for creativity high enough that current AI can't reach it, I suspect quite a lot of human works also fail, e.g. that Pratchett's Strata is obviously Ringworld, and that you would exclude from copyright all parts of The Lion King that are based on Hamlet.

> I would still oppose the use of the natural vs. unnatural distinction as the basis of that argument, though.

I'm not sure what you're saying when you "oppose" this. Does that mean you accept that, in principle, there could be some AI which would deserve rights in the category currently (but in principle inaccurately) called "human rights"?

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#296
post #295

Earlier quoted context omitted.

> you may note that "human rights" is a recent invention and not actually enforced worldwide even today. Consider that countries known for stronger interpretation of human rights and freedoms, including intellectual property rights, are also the countries at the forefront of innovation, including technical innovation that laid the foundation for LLMs in the first place. I think that is not a coincidence, and we shoul…

> Consider that countries known for stronger interpretation of human rights and freedoms, including intellectual property rights, are also the countries at the forefront of innovation, including technical innovation that laid the foundation for LLMs in the first place. I think that is not a coincidence, and we should keep it in mind when there is a push to be dismissive of these concepts (which predominantly serves t…

> China is doing pretty well on the innovation front, including with AI.

From transistors to transformers, most of it builds on foundation that comes guess from where. The innovative layer you speak of is fairly thin.

> Correlation does not imply causation.

I give you that. However, without being able to re-run history, correlation is all we have.

> I would even suggest that it is the disregarding of IP rights that enables much of the web

I would suggest that much of the tech that powers the Web, including probably the most popular operating system on which most servers run, is enabled by copyleft, and copyleft cannot exist without the ability to defend it granted by IP rights, the very concept under fire.

> the point of copyright is to encourage new works

I agree on this.

> All this becomes somewhat irrelevant when AI can also create new works

I don’t agree with a phrase “AI can create new works” for reasons such as 1) “AI” is a meaningless term (let it be my revenge for consciousness) or 2) a tool without agency or will should not be X in a sentence “X can Y” (sure, we can maybe on occasion say “hammers can break things”, but if hammers having agency and will was a popular misconception then I would definitely prefer to stick to “hammers can be used to break things”). The “create new works” part is also questionable on a few levels, but it might exceed the scope of this argument.

That aside, I believe lack of copyright enforcement discourages the creation of new works even in presence of these tools, through the mechanism known as “why would I put effort into new work if I don’t effectively own the result”.

> Does that mean you accept that, in principle, there could be some AI which would deserve rights in the category currently (but in principle inaccurately) called "human rights"?

I think if we believe an LLM or some other software is sufficiently close to a human that it deserves human-like rights or just strong abuse protections (cf. octopus in some countries)—without saying whether I believe it possible or not, it really is orthogonal—then we could excuse it reciting some part of Harry Potter in the right context (probably not as work for hire), but it would be moot because we would also be ethically compelled to not subject it to the training and use that enables such recitation in the first place.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#297
post #215

Earlier quoted context omitted.

Based upon legal decisions in the past there is a clear argument that the distinction for fair use is whether a work is substantially different to another. You are allowed to write a book containg information you learned about from another book. There is threshold in academia regarding plagiarism that stands apart from the legal standing. The measure that was used in Gyles v Wilcox was if the new work could substitut…

Training itself involves making infringing copies of protected works. Whether or not inference produces copyrighted material is almost beside the point.

It’s legal if it’s fair use, which is yet decided by court

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#298
post #295

Earlier quoted context omitted.

> Consider that countries known for stronger interpretation of human rights and freedoms, including intellectual property rights, are also the countries at the forefront of innovation, including technical innovation that laid the foundation for LLMs in the first place. I think that is not a coincidence, and we should keep it in mind when there is a push to be dismissive of these concepts (which predominantly serves t…

> China is doing pretty well on the innovation front, including with AI. From transistors to transformers, most of it builds on foundation that comes guess from where. The innovative layer you speak of is fairly thin. > Correlation does not imply causation. I give you that. However, without being able to re-run history, correlation is all we have. > I would even suggest that it is the disregarding of IP rights that e…

> From transistors to transformers, most of it builds on the foundation that comes guess from where. The innovative layer you speak of is fairly thin.

China rises to first place in most cited papers: https://www.science.org/content/article/china-rises-first-pl...

> 1) “AI” is a meaningless term (let it be my revenge for consciousness)

Fair.

But let's say "computer program" in that case. I'm not fussed about definitions.

> 2) a tool without agency or will should not be X in a sentence “X can Y”. Sure, we can maybe on occasion say “hammers can break things”, but if hammers having agency and will was a popular misconception then I would definitely prefer to stick to “hammers can be used to break things”.

Careful.

If you say that humans can only create things with copyright (even if to support copyleft), then the proletariat are the tool that the bourgeois use to create things.

I do not think this is what you intended :P

> That aside, I believe lack of copyright enforcement discourages the creation of new works even in presence of these tools, through the mechanism known as “why would I put effort into new work if I don’t effectively own the result”.

Same reason you commission a work, or even just buy it from a shop: because then you have the thing.

I mean, the cost of getting o3 to create a novel worth of text is about the same as the price of a generic book by an unknown author in a second-hand shop: https://openai.com/api/pricing/

I've not tried o3 yet, but I have tried o1, and as I've said on a different thread today, o1's output is merely OK, not worth publishing as a book — and I don't know how long it will take to get there. But it is displacing blog writers and podcast writers: https://news.ycombinator.com/item?id=44287953

> but it would be moot because we would also be ethically compelled to not subject it to the training and use that enables such recitation in the first place.

Surprising.

While I would seriously consider the possibility it may be unethical to force such an AI to work if it didn't want to, I think giving it the capability, the education, to be capable of making that choice rather than just saying "it doesn't matter if I wanted to or not, I can't", is just education, as per our own.

Still, I think that's coherent. I'm not sure I've fully internalised the implications so I will let it be.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#299

Earlier quoted context omitted.

Copyright doesn't "produce a cultural hellscape." That's just nonsense. Capitalism does because it has editorial control over narratives and their marketing and distribution. Those are completely different phenomena. Removing copyright will not suddenly open the floodgates of creativity because anyone can already create anything. But - and this is the key point - most work is me-too derivative anyway . See for exampl…

Everything is derivative. This boundary you are defending between originality and slop is extremely subjective at best. What harm is slop anyway? If originality is so objectively valuable, then why should its value be systemically enforced? At the intersection of capitalism and copyright, I see a serious problem. Collaboration is encapsulated by competition. Because simple derivative work is illegal, all collaboratio…

We know what the world looks like without copyright and that world has far fewer works created and very few artists who can do it full-time absent patronage or independent wealth.

Banning the nonsense that is character copyright and shortening copyright back down to a reasonable length of time (say, 20 years) would still enable the creation of more culturally-relevant derivative works without pauperizing every artist.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#300
post #298

Earlier quoted context omitted.

> China is doing pretty well on the innovation front, including with AI. From transistors to transformers, most of it builds on foundation that comes guess from where. The innovative layer you speak of is fairly thin. > Correlation does not imply causation. I give you that. However, without being able to re-run history, correlation is all we have. > I would even suggest that it is the disregarding of IP rights that e…

> From transistors to transformers, most of it builds on the foundation that comes guess from where. The innovative layer you speak of is fairly thin. China rises to first place in most cited papers: https://www.science.org/content/article/china-rises-first-pl... > 1) “AI” is a meaningless term (let it be my revenge for consciousness) Fair. But let's say "computer program" in that case. I'm not fussed about definitio…

> https://www.science.org/content/article/china-rises-first-pl...

Per capita?

> If you say that humans can only create things with copyright (even if to support copyleft), then the proletariat are the tool that the bourgeois use to create things.

The wealth gap and the divide is unlikely to be helped if more people are going to be using (and paying for, in whatever way) ML-based tech from a handful of large corporations.

> Same reason you commission a work, or even just buy it from a shop: because then you have the thing.

Simple posession is more about physical necessities. Commissioning or buying artwork from someone is not just about posessing it, it comes with supporting someone financially. I could make icons or basic illustrations for some small project myself, but I would still commission them if I can afford it because that supports an artist who may want some work (as well as building up for more collaborations in future). Here, I would be supporting the opposite of those artists, a thing that was built on those artists’ work without their consent. Some middleman megacorp of the worst kind.

> the cost of getting o3 to

Don’t they operate at a loss for the time being? They will have to make money sooner or later.

> While I would seriously consider the possibility it may be unethical to force such an AI to work if it didn't want to, I think giving it the capability, the education, to be capable of making that choice rather than just saying "it doesn't matter if I wanted to or not, I can't", is just education, as per our own.

This goes way beyond my thought. I assume if we are talking about education it would be a given that running it generating images 24/7 non stop, shutting it down/killing it, etc., is already out of the question.

Post reply on HN