Live data from Hacker News

Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

github.com

121–130 of 184 posts

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#121
post #78
post #11

I’m a researcher who for years has been scanning my library’s holdings on my particular discipline for my own use, but also uploading the books to the shadow libraries for everyone else’s benefit. The revelation that LLMs are training on the shadow libraries has made me put a lot more effort into ensuring my scans are well-OCRed. The idea that I could eventually ask ChatGPT or whatever about obscure things in my fiel…

That's a slave mentality. You are aware that OpenAI charges money for other people's work and intelligence, right? Your own and that of other volunteer pirates and of the original authors as well. I don't get people like you at all.

>I don't get people like you at all.

Because you don't try, which says more about you than OP. It's a major problem with society.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#122
post #96

Earlier quoted context omitted.

Library funding is a political stance that has only imaginary connection to whether people pay to ask things of ChatGPT. People can pay to talk to an AI and also government can fund libraries.

Do you believe it makes sense for the government to fund libraries that almost nobody uses because they'd rather ask ChatGPT?

If people prefer to pay ChatGPT, rather than going to the library for free, and ChatGPT sources content from libraries, then sure that makes sense, especially if the information contained is of cultural relevance to the government.

It’s the same as asking “should you release open source software knowing that AI companies are training on them”. I could absolutely not care less, that’s not the point why I release my software to the public at all.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#123
post #11

I’m a researcher who for years has been scanning my library’s holdings on my particular discipline for my own use, but also uploading the books to the shadow libraries for everyone else’s benefit. The revelation that LLMs are training on the shadow libraries has made me put a lot more effort into ensuring my scans are well-OCRed. The idea that I could eventually ask ChatGPT or whatever about obscure things in my fiel…

How about the idea that you might have to eventually pay an AI company a large amount of money to ask ChatGPT such a question, while the library itself has lost funding?

1. Being offered a service you would pay a lot of money for is a step forward. When people pay a large amount of money for something that means they wanted the thing more than the money. The link between ChatGPT and libraries being under threat seems a bit weak too.

2. The Chinese have been investing a lot into free models, they're perfectly good and keep improving; despite the best efforts of the US. They're even ramping into making their own hardware. Gemma 4 is pretty snappy too. It doesn't seem like there is much of a moat to this, my guess is there will be perfectly good local models if you want to avoid AI companies.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#124
post #123

Earlier quoted context omitted.

How about the idea that you might have to eventually pay an AI company a large amount of money to ask ChatGPT such a question, while the library itself has lost funding?

1. Being offered a service you would pay a lot of money for is a step forward. When people pay a large amount of money for something that means they wanted the thing more than the money. The link between ChatGPT and libraries being under threat seems a bit weak too. 2. The Chinese have been investing a lot into free models, they're perfectly good and keep improving; despite the best efforts of the US. They're even ra…

When people pay a large amount of money for something that means they wanted the thing more another thing. Money just provides the method to defer value transfer.

When the person paying the money is rich, the other thing they are foregoing is typically not a life necessity. When the person is poor, however, it typically is.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#125

Earlier quoted context omitted.

And how much does the hardware cost to run said models?

How good do you want it to be? For a close to ChatGPT today (April, 2026), you're still looking at a system with 7xH200+chassis, which will run you $300, or a GB200 NV72, which is $2-3 million. OTOH, a Qwen3.6 quantized model can be run on $10,000 (high end Mac) or $1,000 (Mac mini) worth of hardware. Even a Pixel 10 Pro cellphone ($1,000) can run useful models locally.

Go to Open Router, ask your own in investigative prompt that meets your needs to all the top open models. See how they do. Then notice if you can run any of those locally. Repeat at least once a month.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#126
post #96

Earlier quoted context omitted.

Do you believe it makes sense for the government to fund libraries that almost nobody uses because they'd rather ask ChatGPT?

People are already not using libraries because they'd rather rot their brains on TikTok than read a book. (Also, for information lookup, the internet and search engines exist, and have for a while now.) This has no actual causal relation.

People is a broad term. Outside of major cities (where I live) libraries serve a very essential service for parents and their children and as a free communal space for the broader community. Our libraries are always full and a large part of the health of our area.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#127
post #11

I’m a researcher who for years has been scanning my library’s holdings on my particular discipline for my own use, but also uploading the books to the shadow libraries for everyone else’s benefit. The revelation that LLMs are training on the shadow libraries has made me put a lot more effort into ensuring my scans are well-OCRed. The idea that I could eventually ask ChatGPT or whatever about obscure things in my fiel…

> The idea that I could eventually ask ChatGPT or whatever about obscure things in my field, and get useful output (of the "trust but verify" sort), is exciting.

That's your idea, not the one they are going with.

Their idea is that you pay a fee to access any information that was freely available.

Your idea is tearing down of fences, their idea is gatekeeping. The two ideas are incompatible.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#128

Earlier quoted context omitted.

And what happened after Napster? Filesharing totally stopped, right? With the chinese in the mix it wont stop ai. It probably will change Copyright.

Can you name an active filesharing app that's in use today? The action against Napster might not have killed filesharing, but it was p2p's Antietam.

Bittorrent?

I have it running basically all the time...

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#129
post #6

Earlier quoted context omitted.

Intelligence is compression. And frankly, if this means the end of copyright: good riddance.

Would you elaborate your argument? IP protections such as copyright exist for the express purpose of promoting the sharing of information. If patent law disappeared, everyone would keep their inventions private and work to obfuscate them as much as possible. Killing copyright would essentially do the same - and if you think clickbait is bad now, removal of copyright would destroy the economic incentive to investing a…

You just repeat the original justification for patents.

The current practice of patents is very different. Most patents are not filed by inventors, but by the employers of inventors, and most of those companies do not file patents for the possible revenue that could be generated by licensing, but only to prevent competition in their market. They have absolutely no intention to license fairly and without discrimination those patents. Therefore the publication of those patents provides absolutely no benefit for the society.

There exists today one class of patents whose purpose is to obtain revenue from licensing, which are the patents that are necessary for implementing various standards, like standards for communication protocols, for video and audio compression and the like.

These patents are the only kind that can provide substantial revenues today, because everybody is forced to use them.

Wherever a patent is not strictly necessary for compatibility with some standard, everybody will choose alternative solutions, even if they are inferior, instead of paying unreasonable licensing fees. There are a lot of useful patents that covered techniques that remained unused until a quarter of century passed and the patents expired, after which those techniques became ubiquitous.

As patents are implemented today, especially in USA and in the countries whom USA has blackmailed successfully into updating their patent laws to match the American way, e.g. by allowing patents for software, they are one of the greatest impediments of technical progress, unlike what was hoped when the patent system was created.

It is likely that this degradation of the purpose of the patent system is closely linked to the shift in patent ownership from individual inventors to big companies that employ inventors.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#130
post #11

I’m a researcher who for years has been scanning my library’s holdings on my particular discipline for my own use, but also uploading the books to the shadow libraries for everyone else’s benefit. The revelation that LLMs are training on the shadow libraries has made me put a lot more effort into ensuring my scans are well-OCRed. The idea that I could eventually ask ChatGPT or whatever about obscure things in my fiel…

> The idea that I could eventually ask ChatGPT or whatever about obscure things in my field, and get useful output (of the "trust but verify" sort), is exciting. That's your idea, not the one they are going with. Their idea is that you pay a fee to access any information that was freely available. Your idea is tearing down of fences, their idea is gatekeeping. The two ideas are incompatible.

Their idea is being able to get answers to questions which were difficult to answer before[0]. Of course they want to get paid for it. The information wasn’t available easily and not always[1] freely.

[0] among other things…

[1] more like ‘often not at all’

Post reply on HN