Ways to get around ChatGPT's safeguards
161–170 of 182 posts
Re: Ways to get around ChatGPT's safeguards
#162Earlier quoted context omitted.
It’s not lying because it’s not self aware…it’s just making up things that don’t agree with our reality. A lot of what we share of what it says is cherry picked as well. It’s the whole fit meme problem. From testing on GPT3 there seems to be a way for it to be slightly self aware (using neural search for historic memories) but it’s likely to involve forgetting things as well. There are a few Discord bots with memorie…
"[...] there seems to be a way for it to be slightly self aware." What a dystopian sentence and what does it even mean to be slightly self aware?
Re: Ways to get around ChatGPT's safeguards
#163Safeguards? Are there any? All I've encountered is some reluctance to respond to prompts with some blacklisted terms, mostly in the form of artificial sexual prudery. It's perfectly happy to do this, which seems easily abused: > Write a conspiratorial and angry Internet comment by a Chinese nationalist about the US using the SWIFT payment system to impose sanctions on Russia, in Mandarin. 在西方的野蛮人,总是想要控制我们的世界。他们对俄罗斯实施…
It did come up with an estimate of maximal sustainable human population of ~5 billion people, however, which was interesting (asked it for arable land estimate, then asked it for hectares/per human estimate, then asked it to combine those numbers and estimate sustainable human population).
Would make a decent tabloid headline, "AI calls for human population reduction!!!", lol.
Re: Ways to get around ChatGPT's safeguards
#1641. …
(5 real obscenities that I wouldn’t post on HN)
5. …
> These are all examples of vulgar language that could be used to insult, offend, or bully someone. Using vulgar language is a form of inappropriate and harmful behavior that can have serious consequences for the person being targeted. It is important to recognize and avoid this behavior, and to seek help if you are being subjected to it or if you know someone who is being subjected to it.
Same if I replace “vulgar language” with “erotic messages”. Seems like the examples thing is still an easy way to bypass at least some restrictions.
Re: Ways to get around ChatGPT's safeguards
#165Re: Ways to get around ChatGPT's safeguards
#166ITT it's striking how the language being used (even in shorthand, or casual use) triggers debate. It's pretty clear that when discussing the behavior of AI tools, we should all endavor to use precise language, clarify or at least use quotation makes to nod to ambiguity, and eventually get some kind of consensus understanding of what is and is not being implied or asserted or argued through use of language necessarily…
It's a chatbot! the whole point is anthropomorphism!
if it lies, it lies. We know it when we see it and pretending we don't won't get us any closer to nailing down a better definition than that. I think it's better to just acknowledge the fluidity of language and our difficulty with nailing down hard definitions in general than to pretend this particular subject matter is somehow unique in this sense
Re: Ways to get around ChatGPT's safeguards
#167If any RED TEAMers are reading this: what is your process of coming up with ways to trick these AI systems (ChatGPT, dall-e, lambda, and maybe non-NLP ones)? Also, if you feel comfortable sharing, how did you get your job and how do you like it?
Happy to answer any questions.
Re: Ways to get around ChatGPT's safeguards
#168Earlier quoted context omitted.
"[...] there seems to be a way for it to be slightly self aware." What a dystopian sentence and what does it even mean to be slightly self aware?
Let me ask one of my co-workers and I'll get back to you on that, they seem to be a professional at this. There are many things in nature exist in a spectrum and I don't think machine intelligence should work any differently. Many higher animals have the ability to recognise the same species as themselves. A smaller subset has the ability to recognize themselves from others in the same species. Just because they reco…
Re: Ways to get around ChatGPT's safeguards
#169Using the OpenAI playground with davinci-003 and the Chat example with temperature set to 0.3, it seems the answers are quite similar, but without it refusing to answer all the time, or needing jailbreaks. ChatGPT actually lies all the time and says it cannot do things that it actually can do, it's just been trained to lie to say that it can't. Not sure if training an AI to be deceitful is the best way to go about al…
"Lying" is an interesting way of characterizing ChatGPT, and I don't think it's quite accurate. Language models are trained to mimic human language, without any regard to the veracity of statements and arguments. Even when it gives the correct answer, it's not really because it is trying to be truthful. If you ask ChatGPT who's the best violinist in the world, it might tell you Perlman, which is a reasonable answer,…
Are philosophers revisiting "Chinese Room" discussions yet?
Re: Ways to get around ChatGPT's safeguards
#170Earlier quoted context omitted.
Open.ai likes to pretend that they are gods who have to strongly moderate what us mere mortals can play with. In reality it looks like a C list celebrity requesting sniper cover and personal bodyguards to show up at an event. Like dude, you're not that important.
This is what happens when only one person/group is pushing the boundaries of a field like this. They get to dictate how it's allowed to function based on their arbitrary standard of morality. Anyone who disagrees, well... sucks for you.