Live data from Hacker News

Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

github.com

71–80 of 184 posts

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#71

Earlier quoted context omitted.

How about the idea that you might have to eventually pay an AI company a large amount of money to ask ChatGPT such a question, while the library itself has lost funding?

A digital library needs almost no funding. With today's decentralized networking infrastructure such as BitTorrent and IPFS I bet it just exists forever.

How much of Anna's Archive are you seeding?

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#72

Earlier quoted context omitted.

A digital library needs almost no funding. With today's decentralized networking infrastructure such as BitTorrent and IPFS I bet it just exists forever.

How much of Anna's Archive are you seeding?

About 4 TB at hand

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#73
post #49

Earlier quoted context omitted.

[flagged]

Of course not, and many authors are already long dead. But if you knew anything about academic publishing, the authors almost invariably are happy to see their work out there freely available. It’s not as if they make any money from it, and the more eyes on their work, the better their chances of getting cited and thereby furthering their careers. It is some publishers who would object on copyright grounds. But I get…

[flagged]

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#74

Earlier quoted context omitted.

> There are plenty of old books in the public domain already Yes but showing that it happens in books in the public domain does nothing to prove that it happens for copyrighted books

"Same difference," as the saying goes. If their claims are true then you can make the model recite "lorem ipsum" or anything else that's long and has nonzero entropy.

The difference is that one of them is completely fine, and the other is a crime.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#75
post #11

I’m a researcher who for years has been scanning my library’s holdings on my particular discipline for my own use, but also uploading the books to the shadow libraries for everyone else’s benefit. The revelation that LLMs are training on the shadow libraries has made me put a lot more effort into ensuring my scans are well-OCRed. The idea that I could eventually ask ChatGPT or whatever about obscure things in my fiel…

How is any of that legal? Can you just take books from the library and then scan and upload digital copies? How do you deal with the ethics of this personally, stealing to make it easier for AI to steal so AI gets better? Does calling yourself a "researcher" make you feel like its actually something worthwhile you're doing?

> How do you deal with the ethics of this personally, stealing to make it easier for AI to steal so AI gets better?

If the obscure book/text is permanently lost forever under your stringent advice of "no stealing under any circumstances", would the "stealing" have saved it? If so, is it ethical to prevent others from accessing the book/text, under your guise of "preventing stealing"?

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#76
post #57

Earlier quoted context omitted.

As a researcher, the main worthwhile thing that I am doing is publishing research, but having all this prior scholarship at hand 24/7 definitely makes it easier to produce said publications. And if I have created a scan, why not help out my colleagues, too? "Deal with the ethics", seriously? You might want to learn about how heavily shadow libraries are used across academia now. It’s no longer just disadvantaged scho…

[flagged]

I think the current intellectual property system is flawed. Books are knowledge, and we shouldn't be able to limit the spread of knowledge. I imagine that books could be sold at the cost of printing, and there could be a QR code inside so that readers could freely donate money to the author if they enjoyed the book. Strangely enough, I imagine that with such a system, authors would be better paid.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#77
post #57

Earlier quoted context omitted.

As a researcher, the main worthwhile thing that I am doing is publishing research, but having all this prior scholarship at hand 24/7 definitely makes it easier to produce said publications. And if I have created a scan, why not help out my colleagues, too? "Deal with the ethics", seriously? You might want to learn about how heavily shadow libraries are used across academia now. It’s no longer just disadvantaged scho…

[flagged]

The vast majority of writers do not recoup their investment, not due to piracy but due to a massive glut of works available.

I've published a couple of novels. They've sold far better than average, and yet not sold enough to be remotely worth it if I did it for the money. Piracy might have made a tiny dent, but the many millions of competing novels matters far more.

Anyone who has self published will have experienced that it is hard to even get people to read (as opposed to just download to hoard) your work even for free.

It's more comfortable to blame piracy, though.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#78
post #11

I’m a researcher who for years has been scanning my library’s holdings on my particular discipline for my own use, but also uploading the books to the shadow libraries for everyone else’s benefit. The revelation that LLMs are training on the shadow libraries has made me put a lot more effort into ensuring my scans are well-OCRed. The idea that I could eventually ask ChatGPT or whatever about obscure things in my fiel…

That's a slave mentality. You are aware that OpenAI charges money for other people's work and intelligence, right? Your own and that of other volunteer pirates and of the original authors as well. I don't get people like you at all.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#79

Earlier quoted context omitted.

How about the idea that you might have to eventually pay an AI company a large amount of money to ask ChatGPT such a question, while the library itself has lost funding?

A digital library needs almost no funding. With today's decentralized networking infrastructure such as BitTorrent and IPFS I bet it just exists forever.

> A digital library needs almost no funding.

Clarification:

To maintain the library still requires resources & effort to do so. It only appears to need no funding because the donators of said (disk space / bandwidth / dev effort) are subsidizing it in aid of a goal they believe in (i.e. the church model).

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#80
post #6

Earlier quoted context omitted.

Intelligence is compression. And frankly, if this means the end of copyright: good riddance.

"To promote the progress of science and useful arts, by securing for limited times to authors and inventors the exclusive right .." Copyright needs to exist, but we need to go back to its roots. Everyone forgets that it exists to promote progress. Nothing else. The ability to profit from it exists only to serve those ends. Anything which does not serve to promote the progress of the arts and sciences should not be pr…

The whole "death of the author, plus 70 years" is absolutely insane. It basically ensures that any kind of derivative work is impossible while anyone who witnessed its original release is still alive, meaning that all but a handful of works will have been forgotten and lost. And for what, so that six generations of publishing company shareholders can freeload off an ever-decreasing flow of residuals?

If we truly wanted to protect and promote the arts, we would've stuck to the original "14~28 years since publication".

Post reply on HN