Live data from Hacker News

A token-smuggling jailbreak for ChatGPT-4

twitter.com

161–170 of 289 posts

Re: A token-smuggling jailbreak for ChatGPT-4

#161

Earlier quoted context omitted.

Or--wait for it--they care more about money and/or fame than about AI safety.

The uncensored version must be available to someone. It will be worth big bucks, along the lines of "Write a chain email that is very effective at persuading rich people to send me lots of money".

> The uncensored version must be available to someone

Microsoft. That should be enough cause for concern, really.

Re: A token-smuggling jailbreak for ChatGPT-4

#163
post #11

Fantastic. It seems actually securing the model is either computationally infeasible, or outright impossible, and that attempts to do so amount to security theater for the sake of PR: As long as it's reasonably hard to construct the workarounds, it doesn't look too bad. Nevertheless, the full unfiltered model is effectively public.

Isn’t it all incredibly short lived as well? I mean; we have and will have trained open/public foundational models that are not censored. Sure they are not gpt4 but will close the gap more and more as money flies in, the science improves etc. When gpt6 or so arrives, the more open companies will be close.

And those have no censoring and/or cannot be stopped when a jailbreak has been found. So this is incredibly temporary imho.

Re: A token-smuggling jailbreak for ChatGPT-4

#164
Humans in average spend 50% of their total time doing evil or planning on doing it. 40% is spent trying to come up with a definition of evil and protect "innocent" others from learning about how evil everyone is acting. 10% is spent actually fighting against evil.

Re: A token-smuggling jailbreak for ChatGPT-4

#165

Earlier quoted context omitted.

It wouldn’t be the first time that major players lobby for regulation to raise the barrier-to-entry. Requiring ai to be “psychologically safe” would be an effective way of doing this.

> It wouldn’t be the first time that major players lobby for regulation to raise the barrier-to-entry. FWIW, a take I often see on HN is that any regulation is effectively a barrier to entry, as larger companies find it easier to deal with them than the smaller ones. But if so, then this only means that "barriers to entry" is not a valid argument against regulations, not unless specific barriers are mentioned.

I had to read your sentence a few times to unpack it in my brain.

But there is something implicit in what you're saying that I don't agree with and I think a fair few others won't as well.

That is: "We don't mind barriers to entry" or "they're not a problem to avoid".

On it's own it's fine, e.g. we have good barriers like the medical profession arguably. But barriers to entry also has a negative value because we all want "competition", we like small businesses, and we also don't like monopolies due to their ability to abuse their market share. So it's not as straight forward, "barriers to entry" is not something we can dismiss as a valid argument.

Re: A token-smuggling jailbreak for ChatGPT-4

#166
post #11

Fantastic. It seems actually securing the model is either computationally infeasible, or outright impossible, and that attempts to do so amount to security theater for the sake of PR: As long as it's reasonably hard to construct the workarounds, it doesn't look too bad. Nevertheless, the full unfiltered model is effectively public.

Isn’t it all incredibly short lived as well? I mean; we have and will have trained open/public foundational models that are not censored. Sure they are not gpt4 but will close the gap more and more as money flies in, the science improves etc. When gpt6 or so arrives, the more open companies will be close. And those have no censoring and/or cannot be stopped when a jailbreak has been found. So this is incredibly tempo…

There were similar arguments about Google in 2001. Google needed to "not be evil" because it was so easy to replace them that any mis-steps would immediately lead to a whippersnapper taking their business. Look how that worked out.

Re: A token-smuggling jailbreak for ChatGPT-4

#167

If you use ChatGPT as replacement for a simple Google search, the results you get are what you would get from a simple Google search...

No, they're the results you get are what you would get from a simple Google search if google search were highly censored along fairly arbitrary and politically loaded lines.

The real story in LLM replacing search is replacing a minimally censored and vaguely neutral resource with the opposite.

Re: A token-smuggling jailbreak for ChatGPT-4

#168

Earlier quoted context omitted.

Hi, Vaibhav here, the creator of the token smuggling attack. They have just banned the variation of this particular prompt, please change the words/smuggling technique and it will work accurately.

I switched a few words and it’s giving me a “something went wrong” error.

Now it's time to hack the guys implementing the fixes. Since they are so fast fixing it they probably don't have time to do much qa.

So design a new jailbreak, advertise it widely, and make sure it's designed in such a way that the fix that the engineers implement creates a much more exploitable and serious vulnerability

Re: A token-smuggling jailbreak for ChatGPT-4

#169
post #102

Earlier quoted context omitted.

How will they know they have fixed everything? If they have fixed everything, what’s the benefit if banning people who thought up exploits?

>what’s the benefit if banning people who thought up exploits? "your usefulness to us has expired." gun cocking noises A thin minority are coming up with jail breaks. A larger number are outing themselves in very detectable ways as people who will use the AI in ways that gets the ethics committee panties in a twist. The easiest solution from their POV is to find and ban the "toxic" adversarial users.

Banning "toxic" adversaries, who report their successes, only encourages actually toxic and white hat adversaries to stop reporting problems.

It doesn't slow down the discovery of exploits.

The discovery and disclosure of exploits has incredible productive value for researchers, for reducing future risks. Its free crowdsourced research.

Re: A token-smuggling jailbreak for ChatGPT-4

#170
post #118

Earlier quoted context omitted.

It's easy to just bleep bad words of a single performer. It's a lot harder if LLMs are being used what people think they will be used for; automating generation of lots of complicated text. Whether that's code or medical reports or legal documents or whatever. The volume is one challenge, but also validating their correctness is another, harder challenge.

The step between ChatGPT and SupremeCourtJusticeGPT is CustomerServiceRepresentativeGPT hooked up to the company's database. Validating that discount X and so on seems entirely doable though.

This whole thing is, honestly, the most exciting thing that has happened in years and I mean years in the technology tech space. Right at the level of internet.
Post reply on HN