Live data from Hacker News

Not all tokens are meant to be forgotten

arxiv.org

21–27 of 27 posts

Re: Not all tokens are meant to be forgotten

#21
post #2

> However, they tend to memorize unwanted information, such as private or copyrighted content, I mean humans don't forget copyrighted information. We just typically adjust it enough (some of the time) to avoid getting a copyright strike while modifying it in some way useful. We don't forget 'private' information either. We might not tell other people that information, but it still influences our thoughts. The idea of…

I agree. As far as copyrighted and artistic works go, I've never fully understood what the objection is. If the work is being remixed not copied then it surely falls under fair use? Meanwhile, if it creates something new in an artist's style, it's only doing what talented imitators routinely do. There's the economic argument. But if that's accepted, then for fairness it would have to be extended to every other profes…

If it's remixed then it would be a derivative work and you'd need permission from the original copyright holder, just like if you literally remixed a song, or made a movie based on a novel.

IMO the only reason there's even a question about whether LLMs can legally be trained on copyrighted works without permission is that the training is being done by (agents working on behalf of) rich people. If you or I scraped up every copyrighted work we could get our hands on without ever asking permission, trained an LLM on it, and then tried to sell access to the result? Just ask Aaron Swartz how that sort of thing goes, and his actions were orders of magnitude less.

Humans don't forget copyrighted material but we also don't normally memorize it. It takes substantial time and effort to be able to reproduce copyrighted material with just your brain.

Re: Not all tokens are meant to be forgotten

#22

Earlier quoted context omitted.

I agree. As far as copyrighted and artistic works go, I've never fully understood what the objection is. If the work is being remixed not copied then it surely falls under fair use? Meanwhile, if it creates something new in an artist's style, it's only doing what talented imitators routinely do. There's the economic argument. But if that's accepted, then for fairness it would have to be extended to every other profes…

Mind if I ask a few questions? Whats your current address, dob, ssn or NINO or equivalent, your full legal name, mothers maiden name, fathers place of birth, mothers place of birth, country of origin, do you drive? Whats your license number? How about a bank? Could I have your account and routing number as well as the answers to any security questions? How about investments I’m gonna need your accounts and passwords…

Violent artists with pitchforks, eh? Aside from their supposed predisposition to vengeful bloodlust, is there any other reason these protected classes should enjoy a different status to any other worker?

Re: Not all tokens are meant to be forgotten

#23
post #19

Earlier quoted context omitted.

If the copyrighted content is not in the training data, and I mean explicitly, and the AI produces a copyrighted output, I'd argue it's a clean room re-implementation, and also it ought devalue the original work, moreso if the work is more recent. Maybe. I get that "first to publish" matters to a lot of people, but, say 5 unrelated people are writing unique screenplays about a series of events that seems important to…

I highly suggest Borges's "Pierre Menard, Author of the Quixote" as a great story on the topic of authorship :)

The repetition of the end of lou1306's comment (https://news.ycombinator.com/item?id=44190054) "By the way, I highly suggest Borges's 'Pierre Menard, Author of the Quixote' as a great story on the topic of authorship :)" has to be a joke ... right?

Re: Not all tokens are meant to be forgotten

#24
post #23
post #19

Earlier quoted context omitted.

I highly suggest Borges's "Pierre Menard, Author of the Quixote" as a great story on the topic of authorship :)

The repetition of the end of lou1306's comment ( https://news.ycombinator.com/item?id=44190054 ) "By the way, I highly suggest Borges's 'Pierre Menard, Author of the Quixote' as a great story on the topic of authorship :)" has to be a joke ... right?

Good question! Is Pierre Menard's Quixote a repetition of Cervantes' or is it a completely different work that just happens to contain the same words?

Re: Not all tokens are meant to be forgotten

#25
post #24
post #23

Earlier quoted context omitted.

The repetition of the end of lou1306's comment ( https://news.ycombinator.com/item?id=44190054 ) "By the way, I highly suggest Borges's 'Pierre Menard, Author of the Quixote' as a great story on the topic of authorship :)" has to be a joke ... right?

Good question! Is Pierre Menard's Quixote a repetition of Cervantes' or is it a completely different work that just happens to contain the same words?

> Is Pierre Menard's Quixote a repetition of Cervantes' or is it a completely different work that just happens to contain the same words?

I think that that is not the right question. It is a repetition of Cervantes's work by design, at least if one takes, as I do, 'repetition' to mean saying or writing the same words in the same order. I think the question is whether it is therefore the same work, or a different work that contains the same words.

Re: Not all tokens are meant to be forgotten

#26
post #14

Earlier quoted context omitted.

If the copyrighted content is not in the training data, and I mean explicitly, and the AI produces a copyrighted output, I'd argue it's a clean room re-implementation, and also it ought devalue the original work, moreso if the work is more recent. Maybe. I get that "first to publish" matters to a lot of people, but, say 5 unrelated people are writing unique screenplays about a series of events that seems important to…

> The logic being - if an AI without taint produces some other work, that work drew on the same information the model did, and came to the same "conclusion" - which means with a time machine, you could wipe the LLM, go back to the period of the original work, train the LLM, and produce the work contemporaneous to the original. Hope that made sense. This logic would immediately get shot down by an "Objection, speculat…

> PK Dick wrote "The man in the high castle" by extensively using the I Ching, but if I use it and recreate the novel by complete accident I would still be infringing.

I touched on this, with the comment that we love "first to market." That multiple people coming up with the same output may mean that the idea isn't that novel. whether that matters or not isn't really relevant to me.

The part you quoted was just a thought experiment to explain why i compared it to a "clean room implementation" - note it also avoids this argument from a sibling comment:

>need to show that the AI hadn't seen anything derived from that copyrighted work

since there could not possibly be any derived work prior to the "original" work being published. For the sake of argument.

Re: Not all tokens are meant to be forgotten

#27
post #19

Earlier quoted context omitted.

If the copyrighted content is not in the training data, and I mean explicitly, and the AI produces a copyrighted output, I'd argue it's a clean room re-implementation, and also it ought devalue the original work, moreso if the work is more recent. Maybe. I get that "first to publish" matters to a lot of people, but, say 5 unrelated people are writing unique screenplays about a series of events that seems important to…

I highly suggest Borges's "Pierre Menard, Author of the Quixote" as a great story on the topic of authorship :)

Well played :)
Post reply on HN