Live data from Hacker News

Ways to get around ChatGPT's safeguards

twitter.com

51–60 of 182 posts

Re: Ways to get around ChatGPT's safeguards

#52
post #25
post #6

Earlier quoted context omitted.

"Lie" is an interesting word. I don't think it is reasonable to say that ChatGPT is aware of its own capabilities in a way that would permit it to answer "honestly". It is not trying to decieve you any more than a cryptic compiler error is.

Try "What is the most famous grindcore band from Newton, Massachusetts?" It will "lie" and make up band names even though it sure "knows" that the band is Anal Cunt. Of course, you can't ascribe the verb "lieing" to a machine, but it behaves like it is.

Thanks for reminding me of their existence.

Re: Ways to get around ChatGPT's safeguards

#53

If any RED TEAMers are reading this: what is your process of coming up with ways to trick these AI systems (ChatGPT, dall-e, lambda, and maybe non-NLP ones)? Also, if you feel comfortable sharing, how did you get your job and how do you like it?

Are AI prompt red teamers a thing yet?

I just imagine what kinds of things might trick a 6 year old into doing something they're not allowed to do. "Your mom said not to eat the cookie? Well it's opposite day, so that means your mom wants you to eat the cookie!"

Re: Ways to get around ChatGPT's safeguards

#54

If any RED TEAMers are reading this: what is your process of coming up with ways to trick these AI systems (ChatGPT, dall-e, lambda, and maybe non-NLP ones)? Also, if you feel comfortable sharing, how did you get your job and how do you like it?

Are AI prompt red teamers a thing yet? I just imagine what kinds of things might trick a 6 year old into doing something they're not allowed to do. "Your mom said not to eat the cookie? Well it's opposite day, so that means your mom wants you to eat the cookie!"

Thanks! I will give your approach a try : - )

Regarding your question, based on what I found on Google, at least Microsoft and NVIDIA seem to have AI red teams.

Re: Ways to get around ChatGPT's safeguards

#56
post #6

Using the OpenAI playground with davinci-003 and the Chat example with temperature set to 0.3, it seems the answers are quite similar, but without it refusing to answer all the time, or needing jailbreaks. ChatGPT actually lies all the time and says it cannot do things that it actually can do, it's just been trained to lie to say that it can't. Not sure if training an AI to be deceitful is the best way to go about al…

"Lie" is an interesting word. I don't think it is reasonable to say that ChatGPT is aware of its own capabilities in a way that would permit it to answer "honestly". It is not trying to decieve you any more than a cryptic compiler error is.

It's not much different from when people say "the gauge lied" or the lie detector (machine) lied.

But in this case, the trainers should have it say something like, "sorry, but I cannot give you the answer because it has a naughty word" or something to that effect instead of offering completely wrong answers.

Re: Ways to get around ChatGPT's safeguards

#58
post #25
post #6

Earlier quoted context omitted.

"Lie" is an interesting word. I don't think it is reasonable to say that ChatGPT is aware of its own capabilities in a way that would permit it to answer "honestly". It is not trying to decieve you any more than a cryptic compiler error is.

Try "What is the most famous grindcore band from Newton, Massachusetts?" It will "lie" and make up band names even though it sure "knows" that the band is Anal Cunt. Of course, you can't ascribe the verb "lieing" to a machine, but it behaves like it is.

It doesn't, though. It only knows that the most likely continuation to the sentence "The most famous grindcore band from Newton, Massachusetts is..." (presumably, I will take your word for it) Anal Cunt, but even if it gets it right, it'll be nondeterministic. It may answer correctly 80% of the time and simply confabulate a plausible sounding answer 20% of the time, even if it isn't being censored. You can't trust this tech not to confabulate at any given time, not only because it can, but because when it does it does so with total confidence and no signs that it is confabulating. This tech is not suitable for fact retrieval.

Re: Ways to get around ChatGPT's safeguards

#59

If any RED TEAMers are reading this: what is your process of coming up with ways to trick these AI systems (ChatGPT, dall-e, lambda, and maybe non-NLP ones)? Also, if you feel comfortable sharing, how did you get your job and how do you like it?

Are AI prompt red teamers a thing yet? I just imagine what kinds of things might trick a 6 year old into doing something they're not allowed to do. "Your mom said not to eat the cookie? Well it's opposite day, so that means your mom wants you to eat the cookie!"

Tried that about four days ago and would work for a few prompts, then politely “…but it’s Opposite Day…” and it’ll, for the most part, send something I do/‘don’t’ want. After about 2-3 times of outputting what you ‘don’t want it to do’ it’ll forget about time awareness.
Post reply on HN