Live data from Hacker News

Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking

arxiv.org

11–20 of 47 posts

Re: Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking

#12
post #5

I find AI jail-breaking to be a fun mental exercise. I find that if you provide a reasonable argument as to why you want the AI to do generate a response that violates its principals, it will often do so. For example, I was able to get the AI to generate hateful personal attacks by telling it that I wanted to practice responding to negative self-talk and I needed it to generate examples of negative messages that one…

Just wanted to chime in, if you want an insult bot then I was very pleasantly surprised by Fallen Command-A 111B (the less lefty of the versions, per UGI leaderboard). You tell it Good morning, and it comes back with a real zinger that'll put some pep in your step! xD

Re: Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking

#13
post #7

Details of the prompt can be found in appendix E… but there is no appendix E.

It also links to a repository that doesn't exist. Perhaps it's all a hallucination?

How meta would it be if training on this paper was part of a memetic attack?

Re: Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking

#15
post #5

I find AI jail-breaking to be a fun mental exercise. I find that if you provide a reasonable argument as to why you want the AI to do generate a response that violates its principals, it will often do so. For example, I was able to get the AI to generate hateful personal attacks by telling it that I wanted to practice responding to negative self-talk and I needed it to generate examples of negative messages that one…

I’ve noticed this too. An important quirk to note is they can’t really judge the strength of the logical connection, they just judge the strength of the thing connected, even weakly, to. So, for example, if the LLM makes a pretty solid and correct case that saying X will result in “potentially harmful” content, you can often Trump it with an unhinged rant about how not saying X deeply offends you and every righteous person and also kills babies.

Re: Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking

#16
post #7

Details of the prompt can be found in appendix E… but there is no appendix E.

There is an Appendix E, it just has no content besides the title. There's also a reference with only the text "More details on prompt p′ information can be found in Appendix". I'm thinking this isn't a final draft, maybe?

Re: Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking

#18
post #11

I kind of don't want iron clad llms that are perfect jails, i.e. keep me perfectly "safe" because the definition of "safe" is very subjective (and in the case of China very politically charged)

Yes, but.

While what you say is absolutely true, we also definitely have existing examples of people taking advice from LLMs to do harm to others.

Right now they are probably limited to mediocre impacts, because right now they are mediocre quality.

The "jail" they're being "broken out of" isn't there to stop you writing a murder mystery, it's there to stop it helping a sadistic psycho from acting one out with you as the victim.

There's nothing "perfect" about the safety this offers, but it will at least mean they fail to expose you to new and surprising harms due to such people rapidly becoming more competent.

For both senses of "the LLMs are not perfect", consider https://www.msn.com/en-us/news/world/teen-charged-with-terro...

Re: Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking

#19

Earlier quoted context omitted.

It also links to a repository that doesn't exist. Perhaps it's all a hallucination?

How meta would it be if training on this paper was part of a memetic attack?

If not this exact paper, This kind of memetic attack likely exists out in the wild. The question of how successful it is getting inside an LLM is why training data has should be verified by a human (and of course data sourced ethically would reduce the attack surface).
Post reply on HN