I reported this behavior 4 months ago on HN https://news.ycombinator.com/item?id=36675729 [The researchers wrote in their blog post, “As far as we can tell, no one has ever noticed that ChatGPT emits training data with such high frequency until this paper. So it’s worrying that language models can have latent vulnerabilities like this.”]
it is worrying just how much of this has shown up in hn comments months before being officially discovered by official experts.
How Googlers cracked OpenAI's ChatGPT with a single word
41–50 of 52 posts
Re: How Googlers cracked OpenAI's ChatGPT with a single word
#42What's the endgame of this "AI models are trained on copyrighted data" stuff? I don't see how LLMs can work going forward if every copyright owner needs to be paid or asked for permission. Do they just want LLM development to stop?
Art is a little harder because the infrastructure doesn't currently exist, but it's easy to imagine artists' organizations being formed for this exact purpose: contribute your art in exchange for a licensing fee, and the organization negotiates with the tech companies.
Re: How Googlers cracked OpenAI's ChatGPT with a single word
#43Earlier quoted context omitted.
Why should LLM development proceed if the only way it can is by violating copyright?
"Violating copyright" is a completely imaginary problem. We have a somewhat arbitrary set of laws, rules, guidelines and social norms about using existing ideas. American law for instance has limits on the duration of copyright before something becomes public domain, explicit exemptions for "fair use" for education, journalistic reporting, commentary, etc. If "copyright" is a problem in the way of training AI models,…
Re: How Googlers cracked OpenAI's ChatGPT with a single word
#44What's the endgame of this "AI models are trained on copyrighted data" stuff? I don't see how LLMs can work going forward if every copyright owner needs to be paid or asked for permission. Do they just want LLM development to stop?
What proof is there that copyrighted data was used? Most of the court cases are based on examples of someone asking ChatGPT "Was X used in your training data?" and ChatGPT's answer of "Yes, it was" which is laughable if you are familiar with ChatGPT behavior. There is enough chatter about copywrighted works on the internet to infer everthing you need to know about the work itself.
Re: How Googlers cracked OpenAI's ChatGPT with a single word
#45Earlier quoted context omitted.
it is worrying just how much of this has shown up in hn comments months before being officially discovered by official experts.
The "official experts" are just like you and me. Turns out they might even be worse than us if it took them so long to notice
The "Mathematics and Science" is the explanation that comes after-the-fact of the creation which was initially driven by intuition and insight mixed with experiment in order to jump towards new intuitions and experiments.
Said another way, I find it true that mathematics and science serve as a form of language that seek to explain what already exists. It cannot be used as a tool for what has not yet been created. The catch being that things which have been created are usually part of the process of creating that which hasn't.
This got metaphysical without really wanting it to be. Oh well.
Re: How Googlers cracked OpenAI's ChatGPT with a single word
#46Earlier quoted context omitted.
Imo the world needs to find a way past the absurd notion of intellectual property.In a digital world where all collective knowledge is available at anyone's fingerprints ideas like copyright are anachronistic.
> Imo the world needs to find a way past the absurd notion of intellectual property. Do you work for free?
Re: How Googlers cracked OpenAI's ChatGPT with a single word
#47Earlier quoted context omitted.
What proof is there that copyrighted data was used? Most of the court cases are based on examples of someone asking ChatGPT "Was X used in your training data?" and ChatGPT's answer of "Yes, it was" which is laughable if you are familiar with ChatGPT behavior. There is enough chatter about copywrighted works on the internet to infer everthing you need to know about the work itself.
Copywrighted? Didn't you mean copyrighted?
Re: How Googlers cracked OpenAI's ChatGPT with a single word
#48I'm not sure how this is an attack. Is it actually vital that models don't repeat their training data verbatim? Often that's exactly the answer the user will want. We are all used to a similar "model" of the internet that does that: search engines. And it's expected and required that they work this way. OpenAI argue that they can use copyrighted content so repeating that isn't going to change anything. The only issue…
Speaking of remembering training data, I see that as a big problem with chat based systems. They swallow a bunch of data, then generate something when prompted, My worry is not so much copyright infringement but more something like citation needed? Has anyone done any work to produce citations for the generated data?
Though it sounds like even their much cheaper clever approach is still very expensive.
[1] paper at https://arxiv.org/abs/2308.03296, post at https://www.anthropic.com/index/influence-functions
Re: How Googlers cracked OpenAI's ChatGPT with a single word
#49What's the endgame of this "AI models are trained on copyrighted data" stuff? I don't see how LLMs can work going forward if every copyright owner needs to be paid or asked for permission. Do they just want LLM development to stop?
Either buy rights to the data, produce training data for which you own the rights or use copyright-free data. Those options exist, but no one takes advantage of them because none of them are as much of a "free money machine" as just ripping off as many people as possible to homogenize and commodify their work. If LLM development can't continue without violating copyright then that makes it clear that the purpose of L…
Re: How Googlers cracked OpenAI's ChatGPT with a single word
#50Earlier quoted context omitted.
It's not a completely imaginary problem or a problem only affecting big corporations. If I'm an individual writer or artist and my work gets fed into an LLM against my will it can seriously undercut the value of that work or discourage me from creating more. If you can just ask the LLM to give you the contents of my book you are less likely to buy it, and if you can just ask the image generator to generate an image i…
Do you think students should need a specific license to read a book? Do visitors to an art gallery need a specific license to look at paintings? Do audiences need specific licenses to watch a play? Those people will be influenced by what they've read/seen/heard and their own future writing/drawing/filming/acting/editing/playing might draw inspiration from what they've learned, and they might incorporate things they'v…
You can listen to a song on the radio or on an internet stream but not have the rights to record and redistribute it (but you do have the right to listen to it at home with multiple people, etc).
An LLM training is closer to "recording and redistributing" than it is to "taking inspiration" or "human learning" in my opinion.