Live data from Hacker News

Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

github.com

81–90 of 184 posts

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#81
post #11

I’m a researcher who for years has been scanning my library’s holdings on my particular discipline for my own use, but also uploading the books to the shadow libraries for everyone else’s benefit. The revelation that LLMs are training on the shadow libraries has made me put a lot more effort into ensuring my scans are well-OCRed. The idea that I could eventually ask ChatGPT or whatever about obscure things in my fiel…

How is any of that legal? Can you just take books from the library and then scan and upload digital copies? How do you deal with the ethics of this personally, stealing to make it easier for AI to steal so AI gets better? Does calling yourself a "researcher" make you feel like its actually something worthwhile you're doing?

> How is any of that legal?

He didn't mention legality. The world is rigged, as you can see by head of state participating in both in running and cover up of history's largest CSE. Watch what people are doing in addition to what they are saying.

I for one am tremendously thankful for TFNA's efforts, since I get access to knowledge that I wouldn't have been able to before.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#82
post #78
post #11

I’m a researcher who for years has been scanning my library’s holdings on my particular discipline for my own use, but also uploading the books to the shadow libraries for everyone else’s benefit. The revelation that LLMs are training on the shadow libraries has made me put a lot more effort into ensuring my scans are well-OCRed. The idea that I could eventually ask ChatGPT or whatever about obscure things in my fiel…

That's a slave mentality. You are aware that OpenAI charges money for other people's work and intelligence, right? Your own and that of other volunteer pirates and of the original authors as well. I don't get people like you at all.

I’ve already posted in this thread about how even if OpenAI charges money for its LLM trained on the literature, that doesn’t change the fact that the literature remains available to everyone through the shadow libraries, and advances in AI mean that one can increasingly work with it locally on one’s own computer.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#83
post #6

Earlier quoted context omitted.

Intelligence is compression. And frankly, if this means the end of copyright: good riddance.

It won't mean the end of copyright, at most it will just shift the balance of power from one set of giant corporations to another. Anthropic (predictably) issued many DMCA takedown requests after the claude code leak. Copyright for me, but not for thee.

they didn’t touch the LLM python conversion though which tells you it’s not that simple.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#84
post #78
post #11

I’m a researcher who for years has been scanning my library’s holdings on my particular discipline for my own use, but also uploading the books to the shadow libraries for everyone else’s benefit. The revelation that LLMs are training on the shadow libraries has made me put a lot more effort into ensuring my scans are well-OCRed. The idea that I could eventually ask ChatGPT or whatever about obscure things in my fiel…

That's a slave mentality. You are aware that OpenAI charges money for other people's work and intelligence, right? Your own and that of other volunteer pirates and of the original authors as well. I don't get people like you at all.

Open weight models exist and are critical to us avoiding a future where you have to pay sama a slice of every engineers salary.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#85
post #11

I’m a researcher who for years has been scanning my library’s holdings on my particular discipline for my own use, but also uploading the books to the shadow libraries for everyone else’s benefit. The revelation that LLMs are training on the shadow libraries has made me put a lot more effort into ensuring my scans are well-OCRed. The idea that I could eventually ask ChatGPT or whatever about obscure things in my fiel…

How is any of that legal? Can you just take books from the library and then scan and upload digital copies? How do you deal with the ethics of this personally, stealing to make it easier for AI to steal so AI gets better? Does calling yourself a "researcher" make you feel like its actually something worthwhile you're doing?

You can't steal information don't be silly. You can just not have permission to copy it. Oh no.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#87
post #34

Earlier quoted context omitted.

The Bittorrent ecosystem is still very much around. I’m a cinephile who has a collection of nearly a thousand films in Blu-Ray image format, and 95% of that is off a tracker that is open even, not private. And Soulseek is still known as the P2P source where you can find all kinds of obscure music.

> The Bittorrent ecosystem is still very much around. The point is: When Napster was around, everyone was running it all the time from their dorm rooms; it was ubiquitous. Now most people run something like Spotify or Netflix instead; piracy is niche, streaming is ubiquitous.

Using Spotify or Netflix as the example of people getting cold to file sharing is odd. People use Spotify and Netflix because piracy is a service problem, and streaming apps made it a lot less friction is get music and video than running LimeWire.

Notably, Spotify did not exist and Netflix did not stream video until long after the Napster suit.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#88
post #6
post #4

Ok we can drop the farce now that it isn’t compression at the core, the anthropomorphic bullshit has done the job it was supposed to - Allow us to centralize the knowledge economy at the cost of IP holders and we get to claim the efficiency gains from centralization as the result of technology and force governments to choose “teh future” (and investments ) over maintaining copyright - a massive value reallocation in…

Intelligence is compression. And frankly, if this means the end of copyright: good riddance.

Would you elaborate your argument? IP protections such as copyright exist for the express purpose of promoting the sharing of information. If patent law disappeared, everyone would keep their inventions private and work to obfuscate them as much as possible.

Killing copyright would essentially do the same - and if you think clickbait is bad now, removal of copyright would destroy the economic incentive to investing any effort into content.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#89

Earlier quoted context omitted.

And what happened after Napster? Filesharing totally stopped, right? With the chinese in the mix it wont stop ai. It probably will change Copyright.

Can you name an active filesharing app that's in use today? The action against Napster might not have killed filesharing, but it was p2p's Antietam.

There are many people sharing many files on usenet. There are few open source projects to automate the downloads.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#90
post #6

Earlier quoted context omitted.

Intelligence is compression. And frankly, if this means the end of copyright: good riddance.

Copyright is what facilitates copyleft. Getting rid of IP protections also rids us of GPL, which gave us a few things including the most popular OS in the world. It’s one thing to reject the specifics of IP laws as currently implementated; it’s another thing to celebrate the dismantling of the entire foundation of open source by for-profit corporate interests who sought to do it for decades.

> Copyright is what facilitates copyleft.

Chesterson's fence. The existence of copyleft is the result of being forced to live within the domain of copyright, not the other way around.

> Getting rid of IP protections also rids us of GPL, which gave us a few things including the most popular OS in the world.

Linux became popular because of the persistent effort of Linus & the Linux community into making the kernel better, not because of copyleft.

> It’s one thing to reject the specifics of IP laws as currently implementated; it’s another thing to celebrate the dismantling of the entire foundation of open source by for-profit corporate interests who sought to do it for decades.

There are similar corporate interests who profit off of hoarding decades-old works so they can charge fees to what should've been in the public domain, under the original durations that should've stayed (28/14 years).

What has resulted from the endless extensions of the original terms has been the societal lobotomization of human creativity, with an untold number of works now being forever lost simply because they were derived from what should've been in the public domain.

When having lived in such a society, and recognizing existing copyright laws as the reason why it is creatively in such a state, the celebration of its destruction should not be treated as illogical.

Post reply on HN