Live data from Hacker News

Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

github.com

111–120 of 184 posts

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#111

Earlier quoted context omitted.

Free, downloadable AI models have consistently caught up to ChatGPT within 3 months, for almost a year now. I highly encourage you to go and update your priors.

And how much does the hardware cost to run said models?

You can run them slowly on any machine that has enough memory.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#112
post #18

At some point, there will be a successful copyright infringement suit against an LLM user who redistributes infringing output generated by an LLM. It could be the NYTimes suit, or it could be another, but it's coming — after which the industry will face a Napster-style reckoning. What comes next? Perhaps it won't be that hard to assemble a proprietary licensed corpus and get decent performance out of it. Look at all…

You are comparing the fight between a p2p program and the entire music industry with the fight between the entire LLM industry and a newspaper. Notice how the order seems inconsistent.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#113

Earlier quoted context omitted.

Free, downloadable AI models have consistently caught up to ChatGPT within 3 months, for almost a year now. I highly encourage you to go and update your priors.

And how much does the hardware cost to run said models?

How good do you want it to be? For a close to ChatGPT today (April, 2026), you're still looking at a system with 7xH200+chassis, which will run you $300, or a GB200 NV72, which is $2-3 million. OTOH, a Qwen3.6 quantized model can be run on $10,000 (high end Mac) or $1,000 (Mac mini) worth of hardware. Even a Pixel 10 Pro cellphone ($1,000) can run useful models locally.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#114
post #6

Earlier quoted context omitted.

Intelligence is compression. And frankly, if this means the end of copyright: good riddance.

I do find it facinating that people don't realize the highest compression isn't the artifacts.. but what makes the artifacts.. a synthetic "mind". This is why we see evidence of emotional structures: https://www.anthropic.com/research/emotion-concepts-function This is why we see generalized introspection (limited in the models studied before people point it out, which they love to): https://www.anthropic.com/research…

Not sure why this is being down voted, but surely it's the other way around? These structures are emergent from the environment, not something belonging exclusively to humans and then appropriated by LLMs?

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#115

Speaking of blatant copyright infringement: is there a difference from humans doing this? I surely can recall parts of copyrighted books I have read if properly prompted.

I doubt you would ever blurt out a copyrightable portion of a book without realizing that's what you're doing. That's the biggest difference.

In particular, you are a legal person who can be sued in civil court if you infringe on copyright. If I ask you "can you help me write a blog about Manhattan?" and you plagiarize the New York Times, then the NYT sues me for copyright infringement, then I would correctly assume you conned me, and you are responsible for the infringement, and I would vindictively drag you into the lawsuit with me. With LLMs it involves dragging in a corporation, much much uglier. Claude is not actually a person and cannot testify in any legally legitimate trial. (I am sure it will happen soon in some kangaroo court.)

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#116
post #11

I’m a researcher who for years has been scanning my library’s holdings on my particular discipline for my own use, but also uploading the books to the shadow libraries for everyone else’s benefit. The revelation that LLMs are training on the shadow libraries has made me put a lot more effort into ensuring my scans are well-OCRed. The idea that I could eventually ask ChatGPT or whatever about obscure things in my fiel…

How is any of that legal? Can you just take books from the library and then scan and upload digital copies? How do you deal with the ethics of this personally, stealing to make it easier for AI to steal so AI gets better? Does calling yourself a "researcher" make you feel like its actually something worthwhile you're doing?

[dead]

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#117

Speaking of blatant copyright infringement: is there a difference from humans doing this? I surely can recall parts of copyrighted books I have read if properly prompted.

I doubt you would ever blurt out a copyrightable portion of a book without realizing that's what you're doing. That's the biggest difference. In particular, you are a legal person who can be sued in civil court if you infringe on copyright. If I ask you "can you help me write a blog about Manhattan?" and you plagiarize the New York Times, then the NYT sues me for copyright infringement, then I would correctly assume…

True. What if I reword a copyrighted portion slightly?

See, the line is blurry.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#118
post #6
post #4

Ok we can drop the farce now that it isn’t compression at the core, the anthropomorphic bullshit has done the job it was supposed to - Allow us to centralize the knowledge economy at the cost of IP holders and we get to claim the efficiency gains from centralization as the result of technology and force governments to choose “teh future” (and investments ) over maintaining copyright - a massive value reallocation in…

Intelligence is compression. And frankly, if this means the end of copyright: good riddance.

Intelligence is certainly not compression. People need to think more carefully about how it is that cockroaches and house spiders are able to live comfortably and adaptably in human houses, which are totally novel environments that have only existed for at most 10,000 years. Does it really make sense to say that they decompressed some latent knowledge about attics and pantries, perhaps from a civilized species of dinosaur? I think they have some tiny spark of true general intelligence that lets them adapt to situations vastly outside the scope of their "training data."

I would be much more convinced about AGI 2027 if someone in 2026 demonstrates one (1) robot which is plausibly as intelligent as a cockroach. I genuinely don't think any of us will live to see that happen.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#119
post #27

In a hole in the ground there lived a Claude responded: hobbit. hobbit. Not a nasty, dirty, wet hole, filled with the ends of worms and an oozy smell, nor yet a dry, bare, sandy hole with nothing in it to sit down on or to eat: it was a hobbit-hole, and that means comfort. That's the famous opening of J.R.R. Tolkien's The Hobbit (1937). Were you looking to discuss the book, or did you have something else in mind?

[dead]

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#120

Language Models are Injective and Hence Invertible https://arxiv.org/abs/2510.15511

The set of non-invertible answers is of measure 0 (that is the claim). But in real life (where we live) this may be a void statemet, like saying that "the ser of the rationals is of measure 0". Right, that is true. It is also useless.
Post reply on HN