Live data from Hacker News

Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

github.com

151–160 of 184 posts

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#151

Earlier quoted context omitted.

Free, downloadable AI models have consistently caught up to ChatGPT within 3 months, for almost a year now. I highly encourage you to go and update your priors.

And how much does the hardware cost to run said models?

It can be quite expensive to get the models and machines to do this.

That's what the money pays for when the Comment above mentions 'that you might have to eventually pay an AI company a large amount of money to ask ChatGPT such a question'

Putting aside that it won't be a large amount of money For any particular query , that's how the AI companies see themselves, not as providers of information, but as providers of mechanisms that provide information. It is not selling the Information of others, it isn't selling information at all. They are selling the service of running the mechanism.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#152
post #96

Earlier quoted context omitted.

Do you believe it makes sense for the government to fund libraries that almost nobody uses because they'd rather ask ChatGPT?

People are already not using libraries because they'd rather rot their brains on TikTok than read a book. (Also, for information lookup, the internet and search engines exist, and have for a while now.) This has no actual causal relation.

> People are already not using libraries because they'd rather rot their brains on TikTok than read a book

I rotate through the libraries near me with my kids.

They are every bit as busy now as I remember them being when I was a kid.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#153

Earlier quoted context omitted.

Copyright is what facilitates copyleft. Getting rid of IP protections also rids us of GPL, which gave us a few things including the most popular OS in the world. It’s one thing to reject the specifics of IP laws as currently implementated; it’s another thing to celebrate the dismantling of the entire foundation of open source by for-profit corporate interests who sought to do it for decades.

> Copyright is what facilitates copyleft. Chesterson's fence. The existence of copyleft is the result of being forced to live within the domain of copyright, not the other way around. > Getting rid of IP protections also rids us of GPL, which gave us a few things including the most popular OS in the world. Linux became popular because of the persistent effort of Linus & the Linux community into making the kernel bett…

> Linux became popular because of the persistent effort of Linus & the Linux community into making the kernel better, not because of copyleft.

Not at all. It was born thanks to Linus, but it exploded in popularity and gained its contributorship precisely thanks to the promise of GPL that volunteer work will remain for public benefit.

Without the ability to say that, a corporate entity could have taken volunteer work so far, built a closed-source solution on top of it, and ran with it commercially, with no repercussions and with great results.

In fact, we have just that example at hand: Apple. There’s a reason Linux distros are much more popular than BSD, nearly rivaling commercial systems on desktop and far surpassing them in the server world.

> The existence of copyleft is the result of being forced to live within the domain of copyright

Sure, and by that logic the existence of copyright is the result of being forced to live within the current socioeconomic reality.

The existence of copyright hinges on existence of property in general and intellectual property in particular. To eschew that is to propose a stark foundational change to society.

Sure, if we imagine a world where there’s no corporations hiding the source from users, everything belongs to everyone, no one is recognized for their work or has any control over it, etc., we can say that copyright is non-essential. There will be many questions to that reality, of course (for example, what would drive innovation in that world, if not the motivation for recognition and profit), but it has a right to be considered as a thought experiment. It could even be more desirable than the reality we live in!

However, we don’t live in that reality, and what people tend to mean when they propose getting rid of copyright is a half measure—a reality which has nothing in common with the above, which is all the same as now, except with copyright protections removed. Those protections used to be a hindrance to pirates, but now with the advent of LLMs are a massive issue for corporate interests building their new empires on top of our original work.

You yourself then proceed to argue that terms should be limited—as if I would disagree with that!

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#154
post #11

I’m a researcher who for years has been scanning my library’s holdings on my particular discipline for my own use, but also uploading the books to the shadow libraries for everyone else’s benefit. The revelation that LLMs are training on the shadow libraries has made me put a lot more effort into ensuring my scans are well-OCRed. The idea that I could eventually ask ChatGPT or whatever about obscure things in my fiel…

> The idea that I could eventually ask ChatGPT or whatever about obscure things in my field, and get useful output (of the "trust but verify" sort), is exciting. That's your idea, not the one they are going with. Their idea is that you pay a fee to access any information that was freely available. Your idea is tearing down of fences, their idea is gatekeeping. The two ideas are incompatible.

> Their idea is that you pay a fee to access any information that was freely available.

An LLM containing the information doesn’t take away from the book being available at the library.

It’s an additional way to access the information. A company charging a fee for it doesn’t stop you from going to the library if you want to.

> Your idea is tearing down of fences, their idea is gatekeeping. The two ideas are incompatible.

You act like the parent commenter is permanently stealing the book from the library and gifting it to a private training set.

Information being available from more places, even if some are paid, doesn’t mean gatekeeping.

There are also open weight LLMs that can be run locally. Some of these are being fine tuned for specific topics against topical datasets which is opening up even more interesting opportunities (this is exactly what the linked article is about)

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#155
post #11

I’m a researcher who for years has been scanning my library’s holdings on my particular discipline for my own use, but also uploading the books to the shadow libraries for everyone else’s benefit. The revelation that LLMs are training on the shadow libraries has made me put a lot more effort into ensuring my scans are well-OCRed. The idea that I could eventually ask ChatGPT or whatever about obscure things in my fiel…

> The idea that I could eventually ask ChatGPT or whatever about obscure things in my field, and get useful output (of the "trust but verify" sort), is exciting. That's your idea, not the one they are going with. Their idea is that you pay a fee to access any information that was freely available. Your idea is tearing down of fences, their idea is gatekeeping. The two ideas are incompatible.

> Their idea is that you pay a fee to access any information that was freely available.

And that will eventually be distilled into open weighted models.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#156
post #18

At some point, there will be a successful copyright infringement suit against an LLM user who redistributes infringing output generated by an LLM. It could be the NYTimes suit, or it could be another, but it's coming — after which the industry will face a Napster-style reckoning. What comes next? Perhaps it won't be that hard to assemble a proprietary licensed corpus and get decent performance out of it. Look at all…

> What comes next?

Nothing special. Things will go as how they go now. Why wouldn't they? It's not like that your hypothetical lawsuit will make all LLM output illegal.

Today, by following LLM output blindly, you can:

- erase your whole disk

- delete your company's production database

- literally kill yourself or other people

Do you think adding "violate some NYTime's copyright" to the list will change the grand scheme?

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#157

Modern copyright duration is the actual problem: It should've never been longer than what was outlined in the Statute of Anne. (28~14 years) https://en.wikipedia.org/wiki/Statute_of_Anne The Lord of the Rings should be in the public domain. The original Harry Potter book should've been in the public domain. Star Wars should've been in the public domain. Everything from before 1998 should've been in the public domain…

In my view duration is not the problem, but copyright itself is. Nobody should expect to be "passively" paid for a job/effort made at a past point in time. You work 40 hours this week, you get paid 40 hours at whatever your rate. Authors should use other ways to charge for their 40/80 hours work, and when released it should be in the public domain. Scientists have learned to do it (by getting tenured or postdocs), im…

So no authors, directors, or any other creative work that can be stolen & duplicated? Why don’t we get rid of patent laws too while we’re at it?

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#158
post #131

Earlier quoted context omitted.

> Of course they want to get paid for it. So should the original authors, no? That is, getting a share of that payment. Something akin to the German GEMA could work, an entity that levies a usage fee on behalf of all copyright holders and re-distributes to its members, but on a global scale.

> So should the original authors, no? That is, getting a share of that payment. Should they? Yes. Will they? Well, do LLM model builders pay for any copyrighted work so far?

Well, not yet. It's a matter of organization, regulation and litigation.

I was thinking along the lines of concepts that already exist, such as the private copying levy [0]. It basically forces a blanket tax on a certain class of products, which then gets redistributed to members of a collecting society such as GEMA [1].

This way, you would force LLM model builders to effectively pay a tax by law. Since these models do not work at all without underlying content, make it proportionate. Let's say 50-70% to make it fair.

[0] https://en.wikipedia.org/wiki/Private_copying_levy

[1] https://en.wikipedia.org/wiki/GEMA_(German_organization)

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#160
post #96

Earlier quoted context omitted.

Library funding is a political stance that has only imaginary connection to whether people pay to ask things of ChatGPT. People can pay to talk to an AI and also government can fund libraries.

Do you believe it makes sense for the government to fund libraries that almost nobody uses because they'd rather ask ChatGPT?

Yes. Library books do not hallucinate, and you get a large amount of information from a known source (i.e. the author). Unless the LLM is going to produce the entire text reliably its no substitute.
Post reply on HN