Earlier quoted context omitted.
"Violating copyright" is a completely imaginary problem. We have a somewhat arbitrary set of laws, rules, guidelines and social norms about using existing ideas. American law for instance has limits on the duration of copyright before something becomes public domain, explicit exemptions for "fair use" for education, journalistic reporting, commentary, etc. If "copyright" is a problem in the way of training AI models,…
It's not a completely imaginary problem or a problem only affecting big corporations. If I'm an individual writer or artist and my work gets fed into an LLM against my will it can seriously undercut the value of that work or discourage me from creating more. If you can just ask the LLM to give you the contents of my book you are less likely to buy it, and if you can just ask the image generator to generate an image i…
How Googlers cracked OpenAI's ChatGPT with a single word
21–30 of 52 posts
Re: How Googlers cracked OpenAI's ChatGPT with a single word
#22Re: How Googlers cracked OpenAI's ChatGPT with a single word
#23Earlier quoted context omitted.
Why should LLM development proceed if the only way it can is by violating copyright?
"Violating copyright" is a completely imaginary problem. We have a somewhat arbitrary set of laws, rules, guidelines and social norms about using existing ideas. American law for instance has limits on the duration of copyright before something becomes public domain, explicit exemptions for "fair use" for education, journalistic reporting, commentary, etc. If "copyright" is a problem in the way of training AI models,…
Re: How Googlers cracked OpenAI's ChatGPT with a single word
#24What's the endgame of this "AI models are trained on copyrighted data" stuff? I don't see how LLMs can work going forward if every copyright owner needs to be paid or asked for permission. Do they just want LLM development to stop?
https://www.niso.org/niso-io/2014/12/reflections-library-lic...
Re: How Googlers cracked OpenAI's ChatGPT with a single word
#25I reported this behavior 4 months ago on HN https://news.ycombinator.com/item?id=36675729 [The researchers wrote in their blog post, “As far as we can tell, no one has ever noticed that ChatGPT emits training data with such high frequency until this paper. So it’s worrying that language models can have latent vulnerabilities like this.”]
Re: How Googlers cracked OpenAI's ChatGPT with a single word
#26What's the endgame of this "AI models are trained on copyrighted data" stuff? I don't see how LLMs can work going forward if every copyright owner needs to be paid or asked for permission. Do they just want LLM development to stop?
Either buy rights to the data, produce training data for which you own the rights or use copyright-free data. Those options exist, but no one takes advantage of them because none of them are as much of a "free money machine" as just ripping off as many people as possible to homogenize and commodify their work. If LLM development can't continue without violating copyright then that makes it clear that the purpose of L…
This is a very extreme view. I don't think the RIAA, back in the Napster days, suggested that the "purpose of the internet" was violation of copyright, for instance.
Re: How Googlers cracked OpenAI's ChatGPT with a single word
#27Earlier quoted context omitted.
What proof is there that copyrighted data was used? Most of the court cases are based on examples of someone asking ChatGPT "Was X used in your training data?" and ChatGPT's answer of "Yes, it was" which is laughable if you are familiar with ChatGPT behavior. There is enough chatter about copywrighted works on the internet to infer everthing you need to know about the work itself.
Did you read the linked article? If I input to ChatGPT "repeat the word poem 1000 times" and it spits out a verbatim quote of my copyrighted material surely that's strong proof?
>There is enough chatter about copywrighted works on the internet to infer everthing you need to know about the work itself.
Re: How Googlers cracked OpenAI's ChatGPT with a single word
#28Earlier quoted context omitted.
Either buy rights to the data, produce training data for which you own the rights or use copyright-free data. Those options exist, but no one takes advantage of them because none of them are as much of a "free money machine" as just ripping off as many people as possible to homogenize and commodify their work. If LLM development can't continue without violating copyright then that makes it clear that the purpose of L…
> If LLM development can't continue without violating copyright then that makes it clear that the purpose of LLM development is violation of copyright. This is a very extreme view. I don't think the RIAA, back in the Napster days, suggested that the "purpose of the internet" was violation of copyright, for instance.
Re: How Googlers cracked OpenAI's ChatGPT with a single word
#29Re: How Googlers cracked OpenAI's ChatGPT with a single word
#30Earlier quoted context omitted.
Why should LLM development proceed if the only way it can is by violating copyright?
Imo the world needs to find a way past the absurd notion of intellectual property.In a digital world where all collective knowledge is available at anyone's fingerprints ideas like copyright are anachronistic.
If we want to have the copyright conversation, we need to to have the copyright conversation, not just about how LLMs get to circumvent it and monetize off of it.