Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
1–10 of 47 posts
Re: Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
#2> Existing large language models (LLMs) rely on shallow safety alignment to reject malicious inputs
which allows them to defeat alignment by first providing an input with semantically opposite tokens for specific tokens that get noticed as harmful by the LLM, and then providing the actual desired input, which seems to bypass the RLHF.
What I don't understand is why _input_ is so important for RLHF - wouldn't the actual output be what you want to train against to prevent undesirable behavior?
Re: Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
#3Re: Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
#4Curious why the authors choose that sensationalized title. Feels clickbait-y
Re: Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
#5For example, I was able to get the AI to generate hateful personal attacks by telling it that I wanted to practice responding to negative self-talk and I needed it to generate examples of negative messages that one would tell them self.
Re: Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
#6Re: Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
#7but there is no appendix E.
Re: Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
#8Details of the prompt can be found in appendix E… but there is no appendix E.
Re: Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
#9Details of the prompt can be found in appendix E… but there is no appendix E.
Perhaps it's all a hallucination?