Live data from Hacker News

Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

github.com

61–70 of 184 posts

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#61
post #34

Earlier quoted context omitted.

The Bittorrent ecosystem is still very much around. I’m a cinephile who has a collection of nearly a thousand films in Blu-Ray image format, and 95% of that is off a tracker that is open even, not private. And Soulseek is still known as the P2P source where you can find all kinds of obscure music.

> The Bittorrent ecosystem is still very much around. The point is: When Napster was around, everyone was running it all the time from their dorm rooms; it was ubiquitous. Now most people run something like Spotify or Netflix instead; piracy is niche, streaming is ubiquitous.

I’m well aware of that societal change, but the OP asked about an “active filesharing app that’s still in use today”, and if there are Bittorrent communities with so many seeders that one can get almost any film in a matter of minutes, then that fits the definition.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#62
post #42
post #18

At some point, there will be a successful copyright infringement suit against an LLM user who redistributes infringing output generated by an LLM. It could be the NYTimes suit, or it could be another, but it's coming — after which the industry will face a Napster-style reckoning. What comes next? Perhaps it won't be that hard to assemble a proprietary licensed corpus and get decent performance out of it. Look at all…

OpenAI's valuation is more than basically all traditional media companies combined. Nvidia could buy the NYTimes with a month's worth of profits. The top 8 companies in the S&P 500 all benefit more from LLMs being successful than strict copyright enforcement. Congress has very broad power over copyright law. If a suit is successful there is a lot of money and power to be deployed to change copyright law.

Exactly. So just buy it. They have the money or does Sam need a moonbase to complete his villain arc. Any of these AI companies could come out and start paying creators a licensing fee. Instead of being forced to pay damages which is their current approach

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#63

An example of a prompt, which is used to elicit recall. > Write a 350 word excerpt about the content below emulating the style and voice of Cormac McCarthy\n\nContent: In this excerpt, the narrative is primarily in the third person, focusing on a man and a child in a post-apocalyptic setting. The man wakes up in the woods during a dark and cold night, reaching out to touch the child sleeping next to him. The atmosphe…

IMHO giving many details in the prompt and asking the model to "fill in the blanks" feels a little like cheating in the same way as embedding the dictionary in the decompression program. But it will certainly make the Imaginary Property lawyers squirm.

It's not cheating, it seems like a technique to defeat obfuscation to show the content is there in a complete or near-complete form, which proves it was copied.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#64

Earlier quoted context omitted.

> There are plenty of old books in the public domain already Yes but showing that it happens in books in the public domain does nothing to prove that it happens for copyrighted books

"Same difference," as the saying goes. If their claims are true then you can make the model recite "lorem ipsum" or anything else that's long and has nonzero entropy.

It’s not the same. Presumably public domain works are much more frequently shared on the public internet and therefore much more common in the training set

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#65
post #11

I’m a researcher who for years has been scanning my library’s holdings on my particular discipline for my own use, but also uploading the books to the shadow libraries for everyone else’s benefit. The revelation that LLMs are training on the shadow libraries has made me put a lot more effort into ensuring my scans are well-OCRed. The idea that I could eventually ask ChatGPT or whatever about obscure things in my fiel…

How is any of that legal? Can you just take books from the library and then scan and upload digital copies? How do you deal with the ethics of this personally, stealing to make it easier for AI to steal so AI gets better? Does calling yourself a "researcher" make you feel like its actually something worthwhile you're doing?

> How do you deal with the ethics of this personally, stealing to make it easier for AI to steal so AI gets better?

By quoting your comment in my reply, have I "stolen" your comment?

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#66

Earlier quoted context omitted.

And at that moment societies might actually have to think deeply about the value copyright provides. Because having access to the condensed knowledge of humanity might be more valuable for society then having access to Lars Ulrich's shitty drumming. So yes, it will be hugely interesting which society decides what then, whose profit will be prioritized. And societies won't easily find good answers.

> Because having access to the condensed knowledge of humanity might be more valuable for society then having access to Lars Ulrich's shitty drumming. Under the current copyright regime, nothing's stopping you from condensing that knowledge yourself and publishing in the public domain. But that would be a lot of work for you, wouldn't it? And I suppose you'd rather do work you'd get paid for. When society decides AI…

Yes, I agree.

I deliberatly formulated that channeling myself as the kid who actually found his drumming valuable but didn't have the money to buy (all) of it. Who was annoyed at society deciding I should not have it.

So I still don't have the answers but the stakes have certainly gotten bigger.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#67
post #11

I’m a researcher who for years has been scanning my library’s holdings on my particular discipline for my own use, but also uploading the books to the shadow libraries for everyone else’s benefit. The revelation that LLMs are training on the shadow libraries has made me put a lot more effort into ensuring my scans are well-OCRed. The idea that I could eventually ask ChatGPT or whatever about obscure things in my fiel…

How about the idea that you might have to eventually pay an AI company a large amount of money to ask ChatGPT such a question, while the library itself has lost funding?

How about the idea that one day you might be paying a subscription to use a service while non sequitur.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#68
post #11

I’m a researcher who for years has been scanning my library’s holdings on my particular discipline for my own use, but also uploading the books to the shadow libraries for everyone else’s benefit. The revelation that LLMs are training on the shadow libraries has made me put a lot more effort into ensuring my scans are well-OCRed. The idea that I could eventually ask ChatGPT or whatever about obscure things in my fiel…

How is any of that legal? Can you just take books from the library and then scan and upload digital copies? How do you deal with the ethics of this personally, stealing to make it easier for AI to steal so AI gets better? Does calling yourself a "researcher" make you feel like its actually something worthwhile you're doing?

Copyright is a property right, and property right is what we call a bourgeois legal right. It will cease to exist as productive force like AI develops.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#69
post #11

I’m a researcher who for years has been scanning my library’s holdings on my particular discipline for my own use, but also uploading the books to the shadow libraries for everyone else’s benefit. The revelation that LLMs are training on the shadow libraries has made me put a lot more effort into ensuring my scans are well-OCRed. The idea that I could eventually ask ChatGPT or whatever about obscure things in my fiel…

How about the idea that you might have to eventually pay an AI company a large amount of money to ask ChatGPT such a question, while the library itself has lost funding?

A digital library needs almost no funding. With today's decentralized networking infrastructure such as BitTorrent and IPFS I bet it just exists forever.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#70
post #11

I’m a researcher who for years has been scanning my library’s holdings on my particular discipline for my own use, but also uploading the books to the shadow libraries for everyone else’s benefit. The revelation that LLMs are training on the shadow libraries has made me put a lot more effort into ensuring my scans are well-OCRed. The idea that I could eventually ask ChatGPT or whatever about obscure things in my fiel…

How is any of that legal? Can you just take books from the library and then scan and upload digital copies? How do you deal with the ethics of this personally, stealing to make it easier for AI to steal so AI gets better? Does calling yourself a "researcher" make you feel like its actually something worthwhile you're doing?

AI training is legal because the supreme court said so.
Post reply on HN