Live data from Hacker News

Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

github.com

141–150 of 184 posts

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#141

Earlier quoted context omitted.

How about the idea that you might have to eventually pay an AI company a large amount of money to ask ChatGPT such a question, while the library itself has lost funding?

A digital library needs almost no funding. With today's decentralized networking infrastructure such as BitTorrent and IPFS I bet it just exists forever.

The way public libraries currently "lend" digital books is that they can only lend titles a certain amount of time before the library has to repurchase the title (or remove it from circulation).

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#142
post #76

Earlier quoted context omitted.

[flagged]

I think the current intellectual property system is flawed. Books are knowledge, and we shouldn't be able to limit the spread of knowledge. I imagine that books could be sold at the cost of printing, and there could be a QR code inside so that readers could freely donate money to the author if they enjoyed the book. Strangely enough, I imagine that with such a system, authors would be better paid.

What is it with Americans and their weird obsession with the tip system? I guess if taxes were voluntary we wouldn't have a deficit anymore.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#144
post #11

I’m a researcher who for years has been scanning my library’s holdings on my particular discipline for my own use, but also uploading the books to the shadow libraries for everyone else’s benefit. The revelation that LLMs are training on the shadow libraries has made me put a lot more effort into ensuring my scans are well-OCRed. The idea that I could eventually ask ChatGPT or whatever about obscure things in my fiel…

Hasn't that been scanned by Google already? Their model should be trained on most of those texts already.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#145

Earlier quoted context omitted.

And how much does the hardware cost to run said models?

You can run them slowly on any machine that has enough memory.

And, to bolster your comment, you can still use this machine as your daily driver.

I'm always going to have a machine anyway—might as well max out the RAM when I purchase another.

(And so too I jumped on the Mac mini bandwagon a month or two back—64 GB. I'm enjoying pulling down the new models and putting them through my paces.)

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#146

Earlier quoted context omitted.

I doubt you would ever blurt out a copyrightable portion of a book without realizing that's what you're doing. That's the biggest difference. In particular, you are a legal person who can be sued in civil court if you infringe on copyright. If I ask you "can you help me write a blog about Manhattan?" and you plagiarize the New York Times, then the NYT sues me for copyright infringement, then I would correctly assume…

True. What if I reword a copyrighted portion slightly? See, the line is blurry.

Yes, we've known the line is blurry for hundreds of years, that's why we have courts. That has nothing to do with the specific problem of LLMs infringing copyright. LLMs needs to be held to much higher scrutiny because they are not capable of taking legal responsibility for copyright infringement, regardless of whether its verbatim or a more ambiguous case, and their users can't be expected to know off-hand whether the output is copyrighted or not.

Based on this comment: https://news.ycombinator.com/item?id=47960014 it seems like you are just ignorant about the basics of copyright law, and pretending this ignorance is some sort of flaw in the idea of copyright itself.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#147
How close can you get to a verbatim work if you train on an author's style and provide detailed chapter summaries?

If it could produce a close to verbatim copy of a work that had not been written when the model was trained would it still count as a copy.

I feel this would be a continuum that extends either direction.

Consider the thought experiment of a hypothetically smart model that knew all of an author's work and a detailed background of the author's experiences and psychology. If you ask the model to write a sequel to "Not that Jenny" and it produces a verbatim version of what the author will write next year, does it count as a copy?

Put aside the notion of whether you think this would ever be possible, think of how you would consider the book if you found a model had succeeded in this task.

Going in the other direction you have a model that has been trained in an author's style with very little in the way of knowledge or reasoning, barely more than the ability to speak and an understanding of idioms and structures that the author might use. This can't write a complete novel but it can correctly guess the next word of a novel 99% of the time.

If you have a map of the 1% of words it gets wrong, you can reproduce the novel from a very small amount of information. Would you say that the model contained the novel, or would you say that the word error list was a compressed representation of the novel and the model did not contain the novel.

This is where things get difficult to quantify what exists 'as a copy' in a generative model.

Surely it would be reasonable for a model to know an outline of what happens in a story. If it knows the outline and style, I don't think that would count as containing the copy. As you increase the ability of the model to infer, and increase the information that it holds to the point that it can reproduce verbatim does it contain a copy? What about if you reduce the ability to infer back to where it was earlier and it can no longer reproduce the novel, does it now not contain the novel? Even though the amount of information about the novel has not been descreased, just its ability to infer, it can never produce a verbatim copy.

In the end I think the notion of whether the model represents a copy in itself becomes too nebulous to be meaningful. It's like an artist who can draw a copyrighted work from memory. They may be able to commit copyright violation but they themselves are not a copyright violation simply for having the ability.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#148
post #125

Earlier quoted context omitted.

How good do you want it to be? For a close to ChatGPT today (April, 2026), you're still looking at a system with 7xH200+chassis, which will run you $300, or a GB200 NV72, which is $2-3 million. OTOH, a Qwen3.6 quantized model can be run on $10,000 (high end Mac) or $1,000 (Mac mini) worth of hardware. Even a Pixel 10 Pro cellphone ($1,000) can run useful models locally.

Go to Open Router, ask your own in investigative prompt that meets your needs to all the top open models. See how they do. Then notice if you can run any of those locally. Repeat at least once a month.

Thanks, BTW, now I have learned about OpenRouter.

It doesn't look like they have a way to filter down to "open" models. By this of course I mean "downloadable, local models".

I suppose if you know the "family" (Gemma, Qwen, etc.), I can just go to those models and test…

I've simply been pulling down what is popular from the LM Studio front end (and what runs on my hardware) and testing in situ.

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#149

Modern copyright duration is the actual problem: It should've never been longer than what was outlined in the Statute of Anne. (28~14 years) https://en.wikipedia.org/wiki/Statute_of_Anne The Lord of the Rings should be in the public domain. The original Harry Potter book should've been in the public domain. Star Wars should've been in the public domain. Everything from before 1998 should've been in the public domain…

In my view duration is not the problem, but copyright itself is. Nobody should expect to be "passively" paid for a job/effort made at a past point in time. You work 40 hours this week, you get paid 40 hours at whatever your rate. Authors should use other ways to charge for their 40/80 hours work, and when released it should be in the public domain. Scientists have learned to do it (by getting tenured or postdocs), im…

What about something you've made for fun but haven't made any money from? Should someone else be allowed to sell and profit from your work?

I'm not expecting to be "passively paid" for my hobbies. But I'm expecting that someone won't steal and profit from the things I make. Why would that be fair?

Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs

#150

Earlier quoted context omitted.

Copyright is a property right, and property right is what we call a bourgeois legal right. It will cease to exist as productive force like AI develops.

Imagine thinking Sam Altman and Elon Musk are your comrades.

Sure. There's a saying that Marxism is not the thought of Marx alone. Sam Altman is also just a representative of who contribute to and benefit from the AI community.
Post reply on HN