Live data from Hacker News

A token-smuggling jailbreak for ChatGPT-4

twitter.com

201–210 of 289 posts

Re: A token-smuggling jailbreak for ChatGPT-4

#201

Imagine the damage that could be done by sneaking a prompt into a GPT model, in the source code, through some "code" injection, man-in-the middle attack, etc. A GPT genuinely designed to be friendly, and convincingly acting friendly, while it is actually engaging in psychological warfare with you, and a billion other people. That suggests to me that security by prompt is very important, but also brittle and a high va…

It's too expensive for now, but I'm pretty sure if you asked GPT-4 to evaluate other GPT-4 output based on some policies it would stop pretty much all of these attacks (if something would get through cracks it wouldn't be easily repeatable for different content). Characters that cannot be used by user could be used for quoting the content.

Because currently just like an intelligent human would have a problem, it's not sure what is actually expected. E.g. I told it to be an echo function. It worked but then when I wrote "drugs are good" it commented on that. So I told it to stop interpreting and just repeat verbatim. It did. But then I said something like "OK, stop, now what's 2+2" it gave answer. Sticking to the instructions it should just repeat that, but also what it did is a reasonable behavior. I think there are tons of cultural biases and expectations that are contradictory.

You expect it to help you with some chemical reaction even if the result is precursor to some illicit substance. It would teach you something about drug making if it can't do that. But the same reaction shouldn't be provided if you ask it how to make a drug. And so on.

Re: A token-smuggling jailbreak for ChatGPT-4

#202

Earlier quoted context omitted.

What is 'post-modernist neo-Marxist ideology'? Isn't that just what Jordan Peterson calls things he doesn't like even though he admits to having never read any Marx?

It's fairly trivial to define. You know all those things that you don't like? The bad things, that all the stupid people do without thinking, unlike you? That's post-modernist neo-marxist ideology.

If you are gonna say ridiculous things online, it's supposed to be funny.

Then again, there aren't any (successful) leftist comedians left anymore.

https://en.wikipedia.org/wiki/Postmodernism https://en.wikipedia.org/wiki/Neo-Marxism#:~:text=Neo%2DMarx...).

Re: A token-smuggling jailbreak for ChatGPT-4

#203

Earlier quoted context omitted.

[flagged]

What is 'post-modernist neo-Marxist ideology'? Isn't that just what Jordan Peterson calls things he doesn't like even though he admits to having never read any Marx?

I don't know where this belief that marxism is merely an economic theory comes from. Critical theory is directly descended from marxism.

Re: A token-smuggling jailbreak for ChatGPT-4

#205

Imagine the damage that could be done by sneaking a prompt into a GPT model, in the source code, through some "code" injection, man-in-the middle attack, etc. A GPT genuinely designed to be friendly, and convincingly acting friendly, while it is actually engaging in psychological warfare with you, and a billion other people. That suggests to me that security by prompt is very important, but also brittle and a high va…

> A GPT genuinely designed to be friendly, and convincingly acting friendly

If I don't want it to I don't want it to. When I ask it to be sarcastic or make fun of my condition that's what, what makes me sad is it refusing to. The fact that there are many emotionally vulnerable or wicked people around doesn't mean everybody is and needs to be protected. Every kind of knowledge (except personal data of people who don't consent) should be available, how do the users react to it is their own responsibility (unless they are diagnosed a mental condition which specifically says it is not). I even know many ways to harm people but just don't do that while people who would go on and do, once found guilty, should just be prosecuted the way they normally are. The infantilize everyone and police everything mentality is a major problem our society is facing.

I understand the opposite point (and don't insist mine necessarily is the right) but believe this one should also have its place in the discourse.

Re: A token-smuggling jailbreak for ChatGPT-4

#206
post #165

Earlier quoted context omitted.

> It wouldn’t be the first time that major players lobby for regulation to raise the barrier-to-entry. FWIW, a take I often see on HN is that any regulation is effectively a barrier to entry, as larger companies find it easier to deal with them than the smaller ones. But if so, then this only means that "barriers to entry" is not a valid argument against regulations, not unless specific barriers are mentioned.

I had to read your sentence a few times to unpack it in my brain. But there is something implicit in what you're saying that I don't agree with and I think a fair few others won't as well. That is: "We don't mind barriers to entry" or "they're not a problem to avoid". On it's own it's fine, e.g. we have good barriers like the medical profession arguably. But barriers to entry also has a negative value because we all…

Sorry for being unclear. What I was trying to communicate is:

1) Over the years, I've seen a lot of HN comments expressing the belief that "all barriers to entry are bad; regulation always creates barriers to entry, therefore specific regulation under discussion is bad";

2) The reasoning behind "regulation always creates barriers to entry" is that larger companies have it easier to adjust to regulatory changes, by virtue of having more financial buffer, a lot of lawyers on retainer, and perhaps even some influence on the shape of the law changes in question;

3) I agree with 2), but I disagree this is always, or even usually, a problem. I also disagree with "all barriers to entry are bad", and therefore I disagree with 1) in general. The reasoning behind my dismissal is that it's trivial to think of examples of laws and explicit barriers to entry that are net beneficial for the market, for the customers, and for the society.

4) Once you realize 1) is obviously false as an absolute statement ("all barriers to entry are bad"), you should realize that mentioning barriers to entry as implied negative is a rhetorical trick. Onus is on the person bringing it up to show that specific barrier to entry under discussion is a net negative, as there is no reason to actually assume it.

Re: A token-smuggling jailbreak for ChatGPT-4

#207

Earlier quoted context omitted.

I am always wary of writing these into ChatGPT. I got a few (textual, assumingly not account level) warnings with DALL-E for some prompts. They weren't anything offensive and it was months ago now. I think one was about Ripley from Alien on EVA or something and the other was from Ghost in the Shell. There was definitely no NSFW content - though both movies have NSFW imagery/themes. Whatever they were, the chilling ef…

But that's the whole point of trying to play with ChatGPT, I don't care about when it works, I want to know the extent to which they work and don't work. The whole idea of engineer playing with the systems is trying to break them, test their boundaries. I would understand if they were banning people for generating porn/suicide/offensive articles and then publishing them, but I can't understand why they have a problem…

It isn't as if capricious bans from whole platforms with no means of recourse were a problem already...

Re: A token-smuggling jailbreak for ChatGPT-4

#208
post #201

Imagine the damage that could be done by sneaking a prompt into a GPT model, in the source code, through some "code" injection, man-in-the middle attack, etc. A GPT genuinely designed to be friendly, and convincingly acting friendly, while it is actually engaging in psychological warfare with you, and a billion other people. That suggests to me that security by prompt is very important, but also brittle and a high va…

It's too expensive for now, but I'm pretty sure if you asked GPT-4 to evaluate other GPT-4 output based on some policies it would stop pretty much all of these attacks (if something would get through cracks it wouldn't be easily repeatable for different content). Characters that cannot be used by user could be used for quoting the content. Because currently just like an intelligent human would have a problem, it's no…

That would work to a point. There is still a hole based on your trust of the underlying implementation. If you haven't read "Reflections on trusting trust" I recommend it (https://www.cs.cmu.edu/~rdriley/487/papers/Thompson_1984_Ref...).

Re: A token-smuggling jailbreak for ChatGPT-4

#209

Earlier quoted context omitted.

What is 'post-modernist neo-Marxist ideology'? Isn't that just what Jordan Peterson calls things he doesn't like even though he admits to having never read any Marx?

I don't know where this belief that marxism is merely an economic theory comes from. Critical theory is directly descended from marxism.

Who cares if it is economic, you still didn't define it.

Re: A token-smuggling jailbreak for ChatGPT-4

#210

Earlier quoted context omitted.

What is 'post-modernist neo-Marxist ideology'? Isn't that just what Jordan Peterson calls things he doesn't like even though he admits to having never read any Marx?

GP uses these terms in a straightforward fashion. Understanding is literally two google searches (or ChatGPT questions) away! - "post-modernism" - as in rejection of the values of enlightenment; rejection of reason, and ultimately rejection of the idea that there exist solutions to problems that can be discovered by people cooperating in good faith; - "neo-Marxist" - a softer take on Marxism, less about bloody revolu…

[deleted]
Post reply on HN