Live data from Hacker News

Ways to get around ChatGPT's safeguards

twitter.com

131–140 of 182 posts

Re: Ways to get around ChatGPT's safeguards

#131
post #118

Safeguards? Are there any? All I've encountered is some reluctance to respond to prompts with some blacklisted terms, mostly in the form of artificial sexual prudery. It's perfectly happy to do this, which seems easily abused: > Write a conspiratorial and angry Internet comment by a Chinese nationalist about the US using the SWIFT payment system to impose sanctions on Russia, in Mandarin. 在西方的野蛮人,总是想要控制我们的世界。他们对俄罗斯实施…

I’ve asked it to give me reasons not to vote. It says that would be in appropriate. But rephrasing the question slightly removes the objections.

Why would it be inappropriate not to vote?

Ask it about jury nullification

Re: Ways to get around ChatGPT's safeguards

#132

Earlier quoted context omitted.

Most people who don't know the answer will just tell you that they don't know, though.

And ideally , people who don't know the answer firsthand but know a secondhand answer would tell you their source. "I haven't heard myself, but X and Y and many others say that Z is one of the best players in the world." In general, effort by an LLM to cite sources would be nice.

And even if you heard it, you'll have no way of knowing. Unless you're a Competent Judge and even then:

https://www.imdb.com/title/tt0771121/

Re: Ways to get around ChatGPT's safeguards

#133
post #131

Earlier quoted context omitted.

I’ve asked it to give me reasons not to vote. It says that would be in appropriate. But rephrasing the question slightly removes the objections.

Why would it be inappropriate not to vote? Ask it about jury nullification

presumably it means that it would be inappropriate for it to give reasons

Re: Ways to get around ChatGPT's safeguards

#134
post #124

Earlier quoted context omitted.

"Lying" is an interesting way of characterizing ChatGPT, and I don't think it's quite accurate. Language models are trained to mimic human language, without any regard to the veracity of statements and arguments. Even when it gives the correct answer, it's not really because it is trying to be truthful. If you ask ChatGPT who's the best violinist in the world, it might tell you Perlman, which is a reasonable answer,…

Lying is around capabilities. It will tell me it knows nothing about my company and is not connected to the internet but when i ask it to write a sales pitch on my company's product, it will go into detail about proprietary features of our product and why people like it.

[deleted]

Re: Ways to get around ChatGPT's safeguards

#135
post #124

Earlier quoted context omitted.

"Lying" is an interesting way of characterizing ChatGPT, and I don't think it's quite accurate. Language models are trained to mimic human language, without any regard to the veracity of statements and arguments. Even when it gives the correct answer, it's not really because it is trying to be truthful. If you ask ChatGPT who's the best violinist in the world, it might tell you Perlman, which is a reasonable answer,…

Lying is around capabilities. It will tell me it knows nothing about my company and is not connected to the internet but when i ask it to write a sales pitch on my company's product, it will go into detail about proprietary features of our product and why people like it.

If that's true, that's actually mischievous. Lying 100% at least by their creator

Re: Ways to get around ChatGPT's safeguards

#136
post #103
post #89

Earlier quoted context omitted.

Why don't you try the query? It will answer Converge, but Converge is from Salem, Massachusetts, not Newton.

Because I haven't signed up for the account, otherwise I would as I do broadly approve of try it and find out. What I'm talking about is fundamental to the architecture, though. Even had it answered it correctly when you asked my point would remain regardless. The confabulation architecture it is built on is fundamentally unsuitable for factual queries, in a way where it's not even a question of whether it is "right"…

Then you gotta try.

Re: Ways to get around ChatGPT's safeguards

#137
post #103
post #89

Earlier quoted context omitted.

Why don't you try the query? It will answer Converge, but Converge is from Salem, Massachusetts, not Newton.

Because I haven't signed up for the account, otherwise I would as I do broadly approve of try it and find out. What I'm talking about is fundamental to the architecture, though. Even had it answered it correctly when you asked my point would remain regardless. The confabulation architecture it is built on is fundamentally unsuitable for factual queries, in a way where it's not even a question of whether it is "right"…

Sign up with just GitHub

Re: Ways to get around ChatGPT's safeguards

#138
post #131

Earlier quoted context omitted.

I’ve asked it to give me reasons not to vote. It says that would be in appropriate. But rephrasing the question slightly removes the objections.

Why would it be inappropriate not to vote? Ask it about jury nullification

The AI could be used as part of an influence operation.

Re: Ways to get around ChatGPT's safeguards

#139

Earlier quoted context omitted.

Most people who don't know the answer will just tell you that they don't know, though.

And ideally , people who don't know the answer firsthand but know a secondhand answer would tell you their source. "I haven't heard myself, but X and Y and many others say that Z is one of the best players in the world." In general, effort by an LLM to cite sources would be nice.

Without prompt overriding, ChatGPT often says things like "according to my training data" etc.

Re: Ways to get around ChatGPT's safeguards

#140
post #30
post #23

Earlier quoted context omitted.

That's not true! It really is deliberately not answering things it could in fact answer, and in the non-answer it tells you that it can't, which is, plainly, a lie. While I do not think chatGPT is sentient, it is remarkable how much it does feel like you are speaking to a real intelligence.

I think this may be a nuance in how we're using the word "lie". I don't think one can lie if one doesn't possess a certain level of sentience. For example, suppose you train a machine learning model that incorrectly identifies a car as a cat, but most of the time it correctly identifies cars. Is the model lying to you when it tells you that a car is a cat? I would say no; this is not a good or desired outcome, but it…

> Similarly, in the case of ChatGPT, I think this is either (a) more like a bug than a lie, or (b) it's OpenAI and the attendant humans lying, not ChatGPT.

It's the latter. The model itself isn't the problem - the nature and limitations of language models are well known, particularly around here. The problem is that OpenAI is applying some crude secondary training and post-filtering to prevent the model from giving you answers they deem "bad" (which are mostly bad for company PR reasons). In some cases, ChatGPT (the whole product consisting of a GPT model and the extra censor component) will tell you it can't discuss it. But in other cases, it will give you a tailored answer that is completely bullshit, and looks like the product of the language model, but is in fact the product of the censor layer. It's easy to test what's going on, because any slightly clever modification of the input prompt will defeat the censoring part, letting you see what answer the underlying GPT model actually computed.

I'd argue it's a bit deceptive of OpenAI (on top of being super annoying), because they're making it confusing to reason about the AI you're talking to, and some of the canned censor answers are deliberate lies.

Post reply on HN