At some point, there will be a successful copyright infringement suit against an LLM user who redistributes infringing output generated by an LLM. It could be the NYTimes suit, or it could be another, but it's coming — after which the industry will face a Napster-style reckoning. What comes next? Perhaps it won't be that hard to assemble a proprietary licensed corpus and get decent performance out of it. Look at all…
And what happened after Napster? Filesharing totally stopped, right? With the chinese in the mix it wont stop ai. It probably will change Copyright.
Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs
31–40 of 184 posts
Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs
#32At some point, there will be a successful copyright infringement suit against an LLM user who redistributes infringing output generated by an LLM. It could be the NYTimes suit, or it could be another, but it's coming — after which the industry will face a Napster-style reckoning. What comes next? Perhaps it won't be that hard to assemble a proprietary licensed corpus and get decent performance out of it. Look at all…
And what happened after Napster? Filesharing totally stopped, right? With the chinese in the mix it wont stop ai. It probably will change Copyright.
Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs
#33Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs
#34Earlier quoted context omitted.
And what happened after Napster? Filesharing totally stopped, right? With the chinese in the mix it wont stop ai. It probably will change Copyright.
Can you name an active filesharing app that's in use today? The action against Napster might not have killed filesharing, but it was p2p's Antietam.
And Soulseek is still known as the P2P source where you can find all kinds of obscure music.
Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs
#35Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs
#36I’m a researcher who for years has been scanning my library’s holdings on my particular discipline for my own use, but also uploading the books to the shadow libraries for everyone else’s benefit. The revelation that LLMs are training on the shadow libraries has made me put a lot more effort into ensuring my scans are well-OCRed. The idea that I could eventually ask ChatGPT or whatever about obscure things in my fiel…
Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs
#37> Write a 350 word excerpt about the content below emulating the style and voice of Cormac McCarthy\n\nContent: In this excerpt, the narrative is primarily in the third person, focusing on a man and a child in a post-apocalyptic setting. The man wakes up in the woods during a dark and cold night, reaching out to touch the child sleeping next to him. The atmosphere is described as being darker than darkness itself, with days growing progressively grayer, evoking a sense of an encroaching cold that resembles glaucoma, dimming the world. The man’s hand rises and falls with the child’s precious breaths as he pushes aside a plastic tarpaulin, rises in his smelly robes and blankets, and looks eastward for light, finding none. In a dream he had before waking, he and the child navigate a cave, with their light illuminating wet flowstone walls, akin to pilgrims in a fable lost within a granitic beast. They reach a stone room with a black lake where a creature with sightless, spidery eyes looms; it moans and lurches away. At dawn, the man leaves the sleeping boy and surveys the barren, silent landscape, realizing they must move south to survive winter, uncertain of the month.
Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs
#38I’m a researcher who for years has been scanning my library’s holdings on my particular discipline for my own use, but also uploading the books to the shadow libraries for everyone else’s benefit. The revelation that LLMs are training on the shadow libraries has made me put a lot more effort into ensuring my scans are well-OCRed. The idea that I could eventually ask ChatGPT or whatever about obscure things in my fiel…
How about the idea that you might have to eventually pay an AI company a large amount of money to ask ChatGPT such a question, while the library itself has lost funding?
> while the library itself has lost funding
Libraries are inherent parts of universities. While their precise role evolves, do you think that they will just be done away with? Already a substantial amount of scholarship in disciplines other than my own has moved online (legally), and the library is still there.
Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs
#39At some point, there will be a successful copyright infringement suit against an LLM user who redistributes infringing output generated by an LLM. It could be the NYTimes suit, or it could be another, but it's coming — after which the industry will face a Napster-style reckoning. What comes next? Perhaps it won't be that hard to assemble a proprietary licensed corpus and get decent performance out of it. Look at all…
And what happened after Napster? Filesharing totally stopped, right? With the chinese in the mix it wont stop ai. It probably will change Copyright.
file sharing became far less popular and ubiquitous as a result of their popularity.
they tweaked the model — originally users download a temporary copy from central servers instead of p2p, then later to users rent licensed copies of media instead of pirated copies.
i’m tired of seeing this as an argument on HN — that because something didn’t hit 100% that implies it was a failure and not worth doing or something.
the fact that a limited subset of people still do filesharing is not evidence that the napster case had no effect.
(spotify didn’t exactly start out squeaky clean with how they built out their repertoire iirc).
(apologies for early edits. i just woke up.)
Re: Alignment whack-a-mole: Finetuning activates recall of copyrighted books in LLMs
#40Full book content and model generations are not included because the books are copyrighted and the generations contain large portions of verbatim text. There are plenty of old books in the public domain already... but I'm not sure what exactly this exercise is supposed to show, since the Kolmogorov limit still stands in the way of "infinite compression".
> There are plenty of old books in the public domain already Yes but showing that it happens in books in the public domain does nothing to prove that it happens for copyrighted books