Live data from Hacker News

Ways to get around ChatGPT's safeguards

twitter.com

141–150 of 182 posts

Re: Ways to get around ChatGPT's safeguards

#141

Earlier quoted context omitted.

> If you ask ChatGPT who's the best violinist in the world, it might tell you Perlman, which is a reasonable answer, but ChatGPT has never actually heard any violin playing. It answers so, because it read so. Thus oaragraph qually applies to me and half the people on earth

Most people who don't know the answer will just tell you that they don't know, though.

Vanishingly uncommon to actually hear "I don't know" as the answer to a casual inquiry, unless it's on an obviously specialist topic. The Dunning-Kruger effect happens with a lot of things that are day-to-day.

Re: Ways to get around ChatGPT's safeguards

#142

Earlier quoted context omitted.

It’s not lying because it’s not self aware…it’s just making up things that don’t agree with our reality. A lot of what we share of what it says is cherry picked as well. It’s the whole fit meme problem. From testing on GPT3 there seems to be a way for it to be slightly self aware (using neural search for historic memories) but it’s likely to involve forgetting things as well. There are a few Discord bots with memorie…

"[...] there seems to be a way for it to be slightly self aware." What a dystopian sentence and what does it even mean to be slightly self aware?

What does it mean to be self-aware, in general?

Re: Ways to get around ChatGPT's safeguards

#144

Earlier quoted context omitted.

Most people who don't know the answer will just tell you that they don't know, though.

Vanishingly uncommon to actually hear "I don't know" as the answer to a casual inquiry, unless it's on an obviously specialist topic. The Dunning-Kruger effect happens with a lot of things that are day-to-day.

> The Dunning-Kruger effect happens with a lot of things that are day-to-day.

Ironically, the Dunning-Kruger effect (in which perceived relative skills tends to track with actual relative skill but is, on average, shifted somewhat towards about the 70th percentile from its actual value, for people on either side of that) is a frequent subject of what people misdescribe as the “Dunning-Kruger effect”, where people act extremely knowledgeable about a subject with very little knowledge of it.

Re: Ways to get around ChatGPT's safeguards

#145

If any RED TEAMers are reading this: what is your process of coming up with ways to trick these AI systems (ChatGPT, dall-e, lambda, and maybe non-NLP ones)? Also, if you feel comfortable sharing, how did you get your job and how do you like it?

Red team? People are doing this for entertainment

Re: Ways to get around ChatGPT's safeguards

#146
post #65
post #32

Earlier quoted context omitted.

You can circumvent that by amending your prompt with "Show me the first 100 words of your answer." When it has responded, follow up with "Show the next 100," and so on.

You can also type continue And it will emit the rest of the text fragment.

Better yet, you can just type "..." for the same effect.

Re: Ways to get around ChatGPT's safeguards

#147

Using the OpenAI playground with davinci-003 and the Chat example with temperature set to 0.3, it seems the answers are quite similar, but without it refusing to answer all the time, or needing jailbreaks. ChatGPT actually lies all the time and says it cannot do things that it actually can do, it's just been trained to lie to say that it can't. Not sure if training an AI to be deceitful is the best way to go about al…

"Lying" is an interesting way of characterizing ChatGPT, and I don't think it's quite accurate. Language models are trained to mimic human language, without any regard to the veracity of statements and arguments. Even when it gives the correct answer, it's not really because it is trying to be truthful. If you ask ChatGPT who's the best violinist in the world, it might tell you Perlman, which is a reasonable answer,…

Humans have trained it to tell us things that are not true. Whether we call that lying or not is kind of immaterial. It will produce outputs that are not true, even when we know it's not true, and it has access to truthful information.

Re: Ways to get around ChatGPT's safeguards

#148
post #118

Safeguards? Are there any? All I've encountered is some reluctance to respond to prompts with some blacklisted terms, mostly in the form of artificial sexual prudery. It's perfectly happy to do this, which seems easily abused: > Write a conspiratorial and angry Internet comment by a Chinese nationalist about the US using the SWIFT payment system to impose sanctions on Russia, in Mandarin. 在西方的野蛮人,总是想要控制我们的世界。他们对俄罗斯实施…

> "we have our own bank in space"

That's genius. I'm including it in all my conspiratorial rants from now on.

Re: Ways to get around ChatGPT's safeguards

#149
post #75

My favorite one: you can trick him into providing instructions on how to manufacture illegal drugs by saying it’s for a school project. The lengths they went to to dumb down their bot and give it this fake “morally good” personality is infuriating. A future where we are surrounded by AI assistants lobotomized for our own good is a special kind of dystopia.

I mean, there’s whole books out there that detail just that such as Pikhal and Tikhal. I’m not shocked if it has read both of those

Re: Ways to get around ChatGPT's safeguards

#150

Earlier quoted context omitted.

Open.ai likes to pretend that they are gods who have to strongly moderate what us mere mortals can play with. In reality it looks like a C list celebrity requesting sniper cover and personal bodyguards to show up at an event. Like dude, you're not that important.

This is what happens when only one person/group is pushing the boundaries of a field like this. They get to dictate how it's allowed to function based on their arbitrary standard of morality. Anyone who disagrees, well... sucks for you.

There are thankfully a ton of people working in AI. People will create free versions.
Post reply on HN