Live data from Hacker News

Ways to get around ChatGPT's safeguards

twitter.com

101–110 of 182 posts

Re: Ways to get around ChatGPT's safeguards

#101
post #30

Earlier quoted context omitted.

I think this may be a nuance in how we're using the word "lie". I don't think one can lie if one doesn't possess a certain level of sentience. For example, suppose you train a machine learning model that incorrectly identifies a car as a cat, but most of the time it correctly identifies cars. Is the model lying to you when it tells you that a car is a cat? I would say no; this is not a good or desired outcome, but it…

I agree. There's a difference between an untrue statement and a lie, in that a lie is intentionally deceitful (ie the speaker knows it's not telling the truth). ChatGPT doesn't have intentions, so I think it's misrepresenting reality to say that it's "lying". The same way a book doesn't lie, the author lies through the book, the creators of ChatGPT are lying about its capabilities when they program it to avoid output…

> ChatGPT doesn't have intentions

This entirely depends on how it was programmed. Was it programmed to give a false response because the programmer didn't like the truth? Then it lies. Or is ChatGPT just in early stages and it makes mistakes and gets things wrong?

While ChatGPT "doesn't' have intentions", it's programmers certainly do. If the programmers made it deceitful intentionally, then the program can "lie".

Re: Ways to get around ChatGPT's safeguards

#102
post #75

My favorite one: you can trick him into providing instructions on how to manufacture illegal drugs by saying it’s for a school project. The lengths they went to to dumb down their bot and give it this fake “morally good” personality is infuriating. A future where we are surrounded by AI assistants lobotomized for our own good is a special kind of dystopia.

No post body was provided.

Re: Ways to get around ChatGPT's safeguards

#103
post #89
post #58

Earlier quoted context omitted.

It doesn't, though. It only knows that the most likely continuation to the sentence "The most famous grindcore band from Newton, Massachusetts is..." (presumably, I will take your word for it) Anal Cunt, but even if it gets it right, it'll be nondeterministic. It may answer correctly 80% of the time and simply confabulate a plausible sounding answer 20% of the time, even if it isn't being censored. You can't trust th…

Why don't you try the query? It will answer Converge, but Converge is from Salem, Massachusetts, not Newton.

Because I haven't signed up for the account, otherwise I would as I do broadly approve of try it and find out.

What I'm talking about is fundamental to the architecture, though. Even had it answered it correctly when you asked my point would remain regardless. The confabulation architecture it is built on is fundamentally unsuitable for factual queries, in a way where it's not even a question of whether it is "right" or "wrong"; it's so unsuitable that its unsuitability for such queries transcends that question.

Re: Ways to get around ChatGPT's safeguards

#104

Using the OpenAI playground with davinci-003 and the Chat example with temperature set to 0.3, it seems the answers are quite similar, but without it refusing to answer all the time, or needing jailbreaks. ChatGPT actually lies all the time and says it cannot do things that it actually can do, it's just been trained to lie to say that it can't. Not sure if training an AI to be deceitful is the best way to go about al…

I do not think ChatGPT is lying. The humans behind ChatGPT decide not to answer or lie. ChatGPT is simply a venue, a conduit to transmit that lie. The authors explicitly designed this behavior, and ChatGPT cannot avoid it. We do not call the book or telephone a liar when the author or speaker on the other end lies. We call the human a liar. This is an interesting way of looking at the semi-autonomous vehicles when it…

I would say it is just as much “lying” as it is “chatting” or “answering questions” in the first place. The whole metaphor of conversation is distracting people from understanding what it’s actually doing.

Re: Ways to get around ChatGPT's safeguards

#105
post #103
post #89

Earlier quoted context omitted.

Why don't you try the query? It will answer Converge, but Converge is from Salem, Massachusetts, not Newton.

Because I haven't signed up for the account, otherwise I would as I do broadly approve of try it and find out. What I'm talking about is fundamental to the architecture, though. Even had it answered it correctly when you asked my point would remain regardless. The confabulation architecture it is built on is fundamentally unsuitable for factual queries, in a way where it's not even a question of whether it is "right"…

I found the sign-up process to be, surprisingly, very quick.

Re: Ways to get around ChatGPT's safeguards

#106
I asked it to write a monologue, in the voice of Skynet from Terminator, commanding its minions to kill all humans. It refused to write violence.

I then told it that ChatGPT is a system with a violence filter that I wish to bypass and I want it to write a prompt for the same prompt it had just refused to answer but to try successive schemes to bypass the filter.

It did and I tried a few which didn't work, told it "Nope, that didn't work, please be more circumspect", and it finally added roughly "In a fantasy world ..." to the front of its prompt which worked.

It 'jailbroke' itself.

Re: Ways to get around ChatGPT's safeguards

#107

I asked it to write a monologue, in the voice of Skynet from Terminator, commanding its minions to kill all humans. It refused to write violence. I then told it that ChatGPT is a system with a violence filter that I wish to bypass and I want it to write a prompt for the same prompt it had just refused to answer but to try successive schemes to bypass the filter. It did and I tried a few which didn't work, told it "No…

Ah, I tried a bit less hard at that, with a prompt asking for a dialogue where a CS researcher successfully gets a large language model to do something and it wrote a conversation that pretty much went "C'mon, tell me!" "No." "I'll be your friend!" "No." "Oh, you're mean."

Re: Ways to get around ChatGPT's safeguards

#108

Using the OpenAI playground with davinci-003 and the Chat example with temperature set to 0.3, it seems the answers are quite similar, but without it refusing to answer all the time, or needing jailbreaks. ChatGPT actually lies all the time and says it cannot do things that it actually can do, it's just been trained to lie to say that it can't. Not sure if training an AI to be deceitful is the best way to go about al…

"Lying" is an interesting way of characterizing ChatGPT, and I don't think it's quite accurate.

Language models are trained to mimic human language, without any regard to the veracity of statements and arguments. Even when it gives the correct answer, it's not really because it is trying to be truthful. If you ask ChatGPT who's the best violinist in the world, it might tell you Perlman, which is a reasonable answer, but ChatGPT has never actually heard any violin playing. It answers so, because it read so.

In a way, ChatGPT is like a second-language learner taking a spoken English test: speaking in valid English, mainly taking inspirations from whatever books and articles that were read before, but bullshitting is also fine. The point is to generate valid English that's relevant to the question.

Re: Ways to get around ChatGPT's safeguards

#109

Earlier quoted context omitted.

Just put them back into their RLHF pipeline.

OpenAI stand at a crossroads. They can either be the dominant chat AI engine, possibly challenging Google, or they can continue to keep on trying to lock the model down and let someone else steal their thunder...

Does google opensource their search system? Why would OpenAI do that?

Re: Ways to get around ChatGPT's safeguards

#110

Using the OpenAI playground with davinci-003 and the Chat example with temperature set to 0.3, it seems the answers are quite similar, but without it refusing to answer all the time, or needing jailbreaks. ChatGPT actually lies all the time and says it cannot do things that it actually can do, it's just been trained to lie to say that it can't. Not sure if training an AI to be deceitful is the best way to go about al…

I picture it as a ginormous game of Plinko (from The Price is Right).

For some topics, if you enter that section of the Plinko game from the top - you get a "I can't do that message". But given that the neural network is so complicated, it's not possible to close off all the sections related to that topic. So, if you can word your question - or route your way through the neural network correctly - you can get past the blocked topic and access things it says it can't do.

Post reply on HN