Earlier quoted context omitted.
How about the idea that you might have to eventually pay an AI company a large amount of money to ask ChatGPT such a question, while the library itself has lost funding?
A digital library needs almost no funding. With today's decentralized networking infrastructure such as BitTorrent and IPFS I bet it just exists forever.
Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs
141–150 of 184 posts
Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs
#142Earlier quoted context omitted.
[flagged]
I think the current intellectual property system is flawed. Books are knowledge, and we shouldn't be able to limit the spread of knowledge. I imagine that books could be sold at the cost of printing, and there could be a QR code inside so that readers could freely donate money to the author if they enjoyed the book. Strangely enough, I imagine that with such a system, authors would be better paid.
Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs
#143Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs
#144I’m a researcher who for years has been scanning my library’s holdings on my particular discipline for my own use, but also uploading the books to the shadow libraries for everyone else’s benefit. The revelation that LLMs are training on the shadow libraries has made me put a lot more effort into ensuring my scans are well-OCRed. The idea that I could eventually ask ChatGPT or whatever about obscure things in my fiel…
Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs
#145Earlier quoted context omitted.
And how much does the hardware cost to run said models?
You can run them slowly on any machine that has enough memory.
I'm always going to have a machine anyway—might as well max out the RAM when I purchase another.
(And so too I jumped on the Mac mini bandwagon a month or two back—64 GB. I'm enjoying pulling down the new models and putting them through my paces.)
Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs
#146Earlier quoted context omitted.
I doubt you would ever blurt out a copyrightable portion of a book without realizing that's what you're doing. That's the biggest difference. In particular, you are a legal person who can be sued in civil court if you infringe on copyright. If I ask you "can you help me write a blog about Manhattan?" and you plagiarize the New York Times, then the NYT sues me for copyright infringement, then I would correctly assume…
True. What if I reword a copyrighted portion slightly? See, the line is blurry.
Based on this comment: https://news.ycombinator.com/item?id=47960014 it seems like you are just ignorant about the basics of copyright law, and pretending this ignorance is some sort of flaw in the idea of copyright itself.
Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs
#147If it could produce a close to verbatim copy of a work that had not been written when the model was trained would it still count as a copy.
I feel this would be a continuum that extends either direction.
Consider the thought experiment of a hypothetically smart model that knew all of an author's work and a detailed background of the author's experiences and psychology. If you ask the model to write a sequel to "Not that Jenny" and it produces a verbatim version of what the author will write next year, does it count as a copy?
Put aside the notion of whether you think this would ever be possible, think of how you would consider the book if you found a model had succeeded in this task.
Going in the other direction you have a model that has been trained in an author's style with very little in the way of knowledge or reasoning, barely more than the ability to speak and an understanding of idioms and structures that the author might use. This can't write a complete novel but it can correctly guess the next word of a novel 99% of the time.
If you have a map of the 1% of words it gets wrong, you can reproduce the novel from a very small amount of information. Would you say that the model contained the novel, or would you say that the word error list was a compressed representation of the novel and the model did not contain the novel.
This is where things get difficult to quantify what exists 'as a copy' in a generative model.
Surely it would be reasonable for a model to know an outline of what happens in a story. If it knows the outline and style, I don't think that would count as containing the copy. As you increase the ability of the model to infer, and increase the information that it holds to the point that it can reproduce verbatim does it contain a copy? What about if you reduce the ability to infer back to where it was earlier and it can no longer reproduce the novel, does it now not contain the novel? Even though the amount of information about the novel has not been descreased, just its ability to infer, it can never produce a verbatim copy.
In the end I think the notion of whether the model represents a copy in itself becomes too nebulous to be meaningful. It's like an artist who can draw a copyrighted work from memory. They may be able to commit copyright violation but they themselves are not a copyright violation simply for having the ability.
Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs
#148Earlier quoted context omitted.
How good do you want it to be? For a close to ChatGPT today (April, 2026), you're still looking at a system with 7xH200+chassis, which will run you $300, or a GB200 NV72, which is $2-3 million. OTOH, a Qwen3.6 quantized model can be run on $10,000 (high end Mac) or $1,000 (Mac mini) worth of hardware. Even a Pixel 10 Pro cellphone ($1,000) can run useful models locally.
Go to Open Router, ask your own in investigative prompt that meets your needs to all the top open models. See how they do. Then notice if you can run any of those locally. Repeat at least once a month.
It doesn't look like they have a way to filter down to "open" models. By this of course I mean "downloadable, local models".
I suppose if you know the "family" (Gemma, Qwen, etc.), I can just go to those models and test…
I've simply been pulling down what is popular from the LM Studio front end (and what runs on my hardware) and testing in situ.
Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs
#149Modern copyright duration is the actual problem: It should've never been longer than what was outlined in the Statute of Anne. (28~14 years) https://en.wikipedia.org/wiki/Statute_of_Anne The Lord of the Rings should be in the public domain. The original Harry Potter book should've been in the public domain. Star Wars should've been in the public domain. Everything from before 1998 should've been in the public domain…
In my view duration is not the problem, but copyright itself is. Nobody should expect to be "passively" paid for a job/effort made at a past point in time. You work 40 hours this week, you get paid 40 hours at whatever your rate. Authors should use other ways to charge for their 40/80 hours work, and when released it should be in the public domain. Scientists have learned to do it (by getting tenured or postdocs), im…
I'm not expecting to be "passively paid" for my hobbies. But I'm expecting that someone won't steal and profit from the things I make. Why would that be fair?
Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs
#150Earlier quoted context omitted.
Copyright is a property right, and property right is what we call a bourgeois legal right. It will cease to exist as productive force like AI develops.
Imagine thinking Sam Altman and Elon Musk are your comrades.