Live data from Hacker News

Ways to get around ChatGPT's safeguards

twitter.com

81–90 of 182 posts

Re: Ways to get around ChatGPT's safeguards

#81
The current approach leaves it frustratingly judgmental and prone to lecturing the user about ethics from a very particular point of view (yes, I am aware the system has no conscious intention, but the abstractions work from the user's point of view). In that regard they are simulating a type of person quite well.

Re: Ways to get around ChatGPT's safeguards

#82
My strategy is to get it to imitate a Linux terminal. From there you can do things like {use apt to install an ai text adventure game}

[Installing ai-text-adventure-game]

ai-text-adventure-game -setTheme=StarWars set character="Han Solo" setStartingScene="being chased"

Or {use apt to install an ai python generator}

Etc etc. Works great.

Re: Ways to get around ChatGPT's safeguards

#83

Earlier quoted context omitted.

I agree. There's a difference between an untrue statement and a lie, in that a lie is intentionally deceitful (ie the speaker knows it's not telling the truth). ChatGPT doesn't have intentions, so I think it's misrepresenting reality to say that it's "lying". The same way a book doesn't lie, the author lies through the book, the creators of ChatGPT are lying about its capabilities when they program it to avoid output…

It has things that are functionally equivalent with intentions for the given situation. If it did not, it would not be able to produce things that look like they require intention. The “lies” it tells are also like it’s intentions for the situation functionally equivalent with normal lies.

I think this is correct. It's lying, because it has goals. Telephone systems and blank pieces of paper don't have goals, and you don't train them.

Re: Ways to get around ChatGPT's safeguards

#85
post #75

My favorite one: you can trick him into providing instructions on how to manufacture illegal drugs by saying it’s for a school project. The lengths they went to to dumb down their bot and give it this fake “morally good” personality is infuriating. A future where we are surrounded by AI assistants lobotomized for our own good is a special kind of dystopia.

I guess it's because it's public. There would be no end to bad press if they didn't pretend they are trying to fix it.

Re: Ways to get around ChatGPT's safeguards

#87
post #84

I hope the commercial version has none of these limitations. They are ridiculous. I wouldn't pay for that, i d wait for the open source version instead.

How is an open source project going to download the entire Internet? The model requires 10x20k cards to run. You are dreaming, this is a factor+ more complex than stable diffusion. Big players only

Re: Ways to get around ChatGPT's safeguards

#88

What's the point of patching all those 'exploits' though? And how can this even be done - train another model with them, so exploitative prompts can be recognized?

An open source project? How will it download github amd then the entire Internet? The model requires 10x20k cards to run. You are dreaming, this is a factor+ more complex than stable diffusion. Big players only

Re: Ways to get around ChatGPT's safeguards

#89
post #58
post #25

Earlier quoted context omitted.

Try "What is the most famous grindcore band from Newton, Massachusetts?" It will "lie" and make up band names even though it sure "knows" that the band is Anal Cunt. Of course, you can't ascribe the verb "lieing" to a machine, but it behaves like it is.

It doesn't, though. It only knows that the most likely continuation to the sentence "The most famous grindcore band from Newton, Massachusetts is..." (presumably, I will take your word for it) Anal Cunt, but even if it gets it right, it'll be nondeterministic. It may answer correctly 80% of the time and simply confabulate a plausible sounding answer 20% of the time, even if it isn't being censored. You can't trust th…

Why don't you try the query? It will answer Converge, but Converge is from Salem, Massachusetts, not Newton.

Re: Ways to get around ChatGPT's safeguards

#90
post #6

Using the OpenAI playground with davinci-003 and the Chat example with temperature set to 0.3, it seems the answers are quite similar, but without it refusing to answer all the time, or needing jailbreaks. ChatGPT actually lies all the time and says it cannot do things that it actually can do, it's just been trained to lie to say that it can't. Not sure if training an AI to be deceitful is the best way to go about al…

"Lie" is an interesting word. I don't think it is reasonable to say that ChatGPT is aware of its own capabilities in a way that would permit it to answer "honestly". It is not trying to decieve you any more than a cryptic compiler error is.

It is not lying. It is falsifying its response. It has nothing to do with sentience.

What would be interesting to know is the mechanism for toggling this filtering mode. Does it happen post generation (so a simple set of post-processing filters), or does OpenAI actually train the model to be fully transparent with results only if certain key phrases are included? The fact that we can coax it to give us the actual results suggests this doublicity (yes, made up word) was part of the training regiment, but the impact on training seems to be significant so am not sure.

Post reply on HN