Live data from Hacker News

Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

understandingai.org

161–170 of 326 posts

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#161
post #111

Earlier quoted context omitted.

Germany does not have something called "fair use," but it does have provisions for uses that are fair . For example your use of the three words to talk about their copyrighted status is perfectly legal in Germany. That somebody wasn't allowed to use them in a specific way in the past doesn't mean that nobody is allowed to use them in any way.

Of course, but „it’s a short quote so you can use it“ is not true (at least in Germany).

To be pedantic, short quotes (as opposed to short copied fragments that are not used as quotes) are explicitly one of the allowed uses (Zitierbefugnis). You can even quote entire works "in an independent scientific work for the purpose of explaining its content"! https://www.gesetze-im-internet.de/englisch_urhg/englisch_ur...

Generally speaking, exceptions to copyright are based on the appropriateness of the amount of copied content for the given allowed use, so the shorter it is, the more likely it is for copying to be permitted. European copyright law isn't much different from fair use in that respect.

Where it does differ is that the allowed uses are more explicitly enumerated. So Meta would have to argue e.g. based on the exception for scientific works specifically, rather than more general principles.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#163

Earlier quoted context omitted.

The main issue on an economical point of view is that copyright is not the framework we need for social justice and everyone florishing by enjoying pre-existing treasures of human heritage and fairly contributing back. There is no morale and justice ground to leverage on when the system is designed to create wealth bottleneck toward a few recipients. Harry Potter is a great piece of artistic work, and it's nice that…

Capitalism is allergic to second-order cybernetics. First-order systems drive outcomes. "Did it make money?" "Did it increase engagement?" "Did it scale?" These are tight, local feedback loops. They work because they close quickly and map directly to incentives. But they also hide a deeper danger: they optimize without questioning what optimization does to the world that contains it. Second-order cybernetics reason a…

This is a brilliant analysis. Thank you.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#164

Earlier quoted context omitted.

Capitalism is allergic to second-order cybernetics. First-order systems drive outcomes. "Did it make money?" "Did it increase engagement?" "Did it scale?" These are tight, local feedback loops. They work because they close quickly and map directly to incentives. But they also hide a deeper danger: they optimize without questioning what optimization does to the world that contains it. Second-order cybernetics reason a…

[flagged]

And this is not Reddit so please don't.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#165
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

> No one is using this as a substitute for buying the book.

You don't get to say that. Copyright protects the author of a work, but does not bind them to enforce it in any instance. Unlike a trademark, a copyright holder does not lose their protection by allowing unlicensed usage.

It is wholly at the copyright holders discretion to decide which usages they allow and which they do not.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#166
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

I think the argument is less about piracy and more that the model(s output) is a derivative work of Harry Potter, and the rights holder should be paid accordingly when it’s reproduced.

Do you personally pay every time you quote copyrighted books or song lyrics?

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#167

Earlier quoted context omitted.

The main issue on an economical point of view is that copyright is not the framework we need for social justice and everyone florishing by enjoying pre-existing treasures of human heritage and fairly contributing back. There is no morale and justice ground to leverage on when the system is designed to create wealth bottleneck toward a few recipients. Harry Potter is a great piece of artistic work, and it's nice that…

Capitalism is allergic to second-order cybernetics. First-order systems drive outcomes. "Did it make money?" "Did it increase engagement?" "Did it scale?" These are tight, local feedback loops. They work because they close quickly and map directly to incentives. But they also hide a deeper danger: they optimize without questioning what optimization does to the world that contains it. Second-order cybernetics reason a…

and as a consequence the fight of AI vs copyright is one of two capitalists fighting each other. it's not about liberating copyright but about shuffling profits around. regardless of who wins that fight society loses.

it conjures up pictures of two dragons fighting each other instead of attacking us, but make no mistake they are only fighting for the right to attack us. whoever wins is coming for us afterwards

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#168
post #156
post #64

As an experiment I searched Google for "harry potter and the sorcerer's stone text": - the first result is a pdf of the full book - the second result is a txt of the full book - the third result is a pdf of the complete harry potter collection - the fourth result is a txt of the full book (hosted on github funny enough) Further down there are similar copies from the internet archive and dozens of other sites. All in…

> let's not pretend that an LLM that autocompletes a couple lines from harry potter with 50% accuracy is some massive new avenue to piracy No one is claiming this. The corporations developing LLMs are doing so by sampling media without their owners' permission and arguing this is protected by US fair use laws, which is incorrect - as the late AI researcher Suchir Balaji explained in this other article: https://suchir…

It's not clear that it's incorrect.

Re: Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book

#169
post #56

Earlier quoted context omitted.

We cannot get ride of it without finding a way to pay the creators that generate copyrighted works. I’m personally more in favor of significantly reducing the length of the copy right. I think 20-30 years is an interesting range. Artist get roughly a career length of time to profit off their creations, but there is much less incentive for major corporations to buy and horde IP.

We barely pay creators as it is for generating copyrighted works. Nearly every copywritten work is available on the internet, for free, right now . And creators are still getting paid, albeit poorly, but that's a constant throughout history.

The thing about creators is that most of them are paid extremely poorly, and some of them get insanely rich. Joanne Rowling has received more money than a reasonable person could use for her wizard books, but millions of bloggers feeding much more data into AI training sets will never see a cent for their work. For starting authors selling books, this can easily be the difference between writing another book or giving up and taking up another job.

At the moment, there's also a huge difference between who does and who doesn't pay. If I put the HP collection on my website, you betcha Joanne Rowling's team is going to try to take it down. However, because OpenAI designed an AI system where content cannot be removed from its knowledge base and because their pockets are lined with cash for lawyers, it's practically free to violate whatever copyright rules it wants.

Post reply on HN