Live data from Hacker News

How Googlers cracked OpenAI's ChatGPT with a single word

sfgate.com

31–40 of 52 posts

Re: How Googlers cracked OpenAI's ChatGPT with a single word

#31

I'm not sure how this is an attack. Is it actually vital that models don't repeat their training data verbatim? Often that's exactly the answer the user will want. We are all used to a similar "model" of the internet that does that: search engines. And it's expected and required that they work this way. OpenAI argue that they can use copyrighted content so repeating that isn't going to change anything. The only issue…

Speaking of remembering training data, I see that as a big problem with chat based systems. They swallow a bunch of data, then generate something when prompted, My worry is not so much copyright infringement but more something like citation needed?

Has anyone done any work to produce citations for the generated data?

Re: How Googlers cracked OpenAI's ChatGPT with a single word

#32
post #19

Earlier quoted context omitted.

"Violating copyright" is a completely imaginary problem. We have a somewhat arbitrary set of laws, rules, guidelines and social norms about using existing ideas. American law for instance has limits on the duration of copyright before something becomes public domain, explicit exemptions for "fair use" for education, journalistic reporting, commentary, etc. If "copyright" is a problem in the way of training AI models,…

It's not a completely imaginary problem or a problem only affecting big corporations. If I'm an individual writer or artist and my work gets fed into an LLM against my will it can seriously undercut the value of that work or discourage me from creating more. If you can just ask the LLM to give you the contents of my book you are less likely to buy it, and if you can just ask the image generator to generate an image i…

Do you think students should need a specific license to read a book? Do visitors to an art gallery need a specific license to look at paintings? Do audiences need specific licenses to watch a play?

Those people will be influenced by what they've read/seen/heard and their own future writing/drawing/filming/acting/editing/playing might draw inspiration from what they've learned, and they might incorporate things they've learned into their own future work.

Literally every book, song and work of art is "violating copyright" on the thousands of other works that the creator learned from while growing up, if we hold the same standard.

Re: How Googlers cracked OpenAI's ChatGPT with a single word

#33
post #9

Earlier quoted context omitted.

Why should LLM development proceed if the only way it can is by violating copyright?

"Violating copyright" is a completely imaginary problem. We have a somewhat arbitrary set of laws, rules, guidelines and social norms about using existing ideas. American law for instance has limits on the duration of copyright before something becomes public domain, explicit exemptions for "fair use" for education, journalistic reporting, commentary, etc. If "copyright" is a problem in the way of training AI models,…

> If "copyright" is a problem in the way of training AI models, then we should all collectively vote for politicians who fix that problem by updating the laws to make the training explicitly allowed. Problem solved.

Yep. The EU has this, as does Singapore, South Korea and Malaysia. A lot of countries have already recognised that it's not a good idea to restrict AI dev because of IP "rights".

Re: How Googlers cracked OpenAI's ChatGPT with a single word

#34
post #13

Earlier quoted context omitted.

Imo the world needs to find a way past the absurd notion of intellectual property.In a digital world where all collective knowledge is available at anyone's fingerprints ideas like copyright are anachronistic.

Sure, I agree with you at a high level. But if the answer is that LLMs get a pass and the rest of us have to deal with DMCA takedown abuse, inaccessible geolocked content, and 7-figure legal penalties for getting caught downloading a $3.99-to-rent movie, then fuck that. If we want to have the copyright conversation, we need to to have the copyright conversation , not just about how LLMs get to circumvent it and monet…

There is no need for a “conversation”. The concept has simply become obsolete and these laws will cease to exist as they are unenforceable.

Re: How Googlers cracked OpenAI's ChatGPT with a single word

#35
post #34

Earlier quoted context omitted.

Sure, I agree with you at a high level. But if the answer is that LLMs get a pass and the rest of us have to deal with DMCA takedown abuse, inaccessible geolocked content, and 7-figure legal penalties for getting caught downloading a $3.99-to-rent movie, then fuck that. If we want to have the copyright conversation, we need to to have the copyright conversation , not just about how LLMs get to circumvent it and monet…

There is no need for a “conversation”. The concept has simply become obsolete and these laws will cease to exist as they are unenforceable.

They won’t cease to exist unless something happens to challenge them and render them unenforceable.

Re: How Googlers cracked OpenAI's ChatGPT with a single word

#36
post #34

Earlier quoted context omitted.

There is no need for a “conversation”. The concept has simply become obsolete and these laws will cease to exist as they are unenforceable.

They won’t cease to exist unless something happens to challenge them and render them unenforceable.

[flagged]

Re: How Googlers cracked OpenAI's ChatGPT with a single word

#37
post #13
post #9

Earlier quoted context omitted.

Why should LLM development proceed if the only way it can is by violating copyright?

Imo the world needs to find a way past the absurd notion of intellectual property.In a digital world where all collective knowledge is available at anyone's fingerprints ideas like copyright are anachronistic.

> Imo the world needs to find a way past the absurd notion of intellectual property.

Do you work for free?

Re: How Googlers cracked OpenAI's ChatGPT with a single word

#38
post #9
post #2

What's the endgame of this "AI models are trained on copyrighted data" stuff? I don't see how LLMs can work going forward if every copyright owner needs to be paid or asked for permission. Do they just want LLM development to stop?

Why should LLM development proceed if the only way it can is by violating copyright?

LLM development will proceed regardless in countries that don't give a damn about our vision of "copyright".

It doesn't matter if you think copyright makes sense or not. In 20 years, some country will have its own giant LLM trained on copyrighted material and use this to boost their competitive advantage and technological power and development, perhaps so much that the advantage will be tremendous, while we'll stay the underdogs because "my copyrights".

Re: How Googlers cracked OpenAI's ChatGPT with a single word

#40
post #6

I reported this behavior 4 months ago on HN https://news.ycombinator.com/item?id=36675729 [The researchers wrote in their blog post, “As far as we can tell, no one has ever noticed that ChatGPT emits training data with such high frequency until this paper. So it’s worrying that language models can have latent vulnerabilities like this.”]

it is worrying just how much of this has shown up in hn comments months before being officially discovered by official experts.

The "official experts" are just like you and me. Turns out they might even be worse than us if it took them so long to notice
Post reply on HN