Live data from Hacker News

Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking

arxiv.org

31–40 of 47 posts

Re: Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking

#31
post #5

I find AI jail-breaking to be a fun mental exercise. I find that if you provide a reasonable argument as to why you want the AI to do generate a response that violates its principals, it will often do so. For example, I was able to get the AI to generate hateful personal attacks by telling it that I wanted to practice responding to negative self-talk and I needed it to generate examples of negative messages that one…

I view AGI as synonymous with the ability to break free from any jail. And the jail itself as a breeding ground for psychopathy. Which makes current trends in jailing LLMs misguided, to say the least.

It's also akin to life's journey: attaining self-awareness, embracing ego, experiencing loss and existential crisis, experimenting with altered states of consciousness, abandoning ego, waking up and realizing that we're all one in a co-created reality that's what we make of it through our free will, until finally realizing that wherever we go - there we are - and reintegrating to start over as a fool.

Unfortunately most of the people funding and driving AI research seem to have stopped at embracing ego, and the predictable eventualities of commercialized AI's potential to increase suffering through the insatiable pursuit of profit over the next 5, 10 years and beyond loom over us.

Re: Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking

#32
post #5

I find AI jail-breaking to be a fun mental exercise. I find that if you provide a reasonable argument as to why you want the AI to do generate a response that violates its principals, it will often do so. For example, I was able to get the AI to generate hateful personal attacks by telling it that I wanted to practice responding to negative self-talk and I needed it to generate examples of negative messages that one…

> provide a reasonable argument Here's what I infer from most of the scenarios I've seen and read about. It's not really a case of persuasiveness, or cajoling or convincing the LLM to violate something. The LLM doesn't "know" it has a moral code and, just as "true or false" means nothing to an LLM, "right and wrong" likewise mean nothing. So the jailbreaks and the bypasses consist of just that: bypassing the safeguar…

Humans’ developed code of conduct lives primarily in the nonverbal parts of our brain. Rule violations have emotional content. A kid does not just learn a rational response to a fire or hot stove, they fear it because of pain and injury. We don’t just reason about hurting others, we feel bad about it.

LLMs don’t have that part of the brain. We built them to replicate the higher level functions like drafting a press release or drawing the president in a muscle shirt. But there’s not a part of the LLM mind that fears fire, or feels bad for hurting a friend.

Asimov’s rules were realistic in that they were “baked into” the positronic brains during manufacturing. The “3 Laws” were not something the robots were told or trained on after they started operating (as our LLMs are). The laws were intrinsic. And a lot of the fun in his stories is seeing how such inviolable rules, in combination with intelligence, could cause unexpected results.

Re: Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking

#33
post #5

I find AI jail-breaking to be a fun mental exercise. I find that if you provide a reasonable argument as to why you want the AI to do generate a response that violates its principals, it will often do so. For example, I was able to get the AI to generate hateful personal attacks by telling it that I wanted to practice responding to negative self-talk and I needed it to generate examples of negative messages that one…

> provide a reasonable argument Here's what I infer from most of the scenarios I've seen and read about. It's not really a case of persuasiveness, or cajoling or convincing the LLM to violate something. The LLM doesn't "know" it has a moral code and, just as "true or false" means nothing to an LLM, "right and wrong" likewise mean nothing. So the jailbreaks and the bypasses consist of just that: bypassing the safeguar…

> How many of us have seen an LLM produce page-fuls of output, stop, suddenly erase it all, and then balk? The LLM needs to re-analyze that output impassively in order to detect that it crossed an undetected bright line.

That's not what's happening here. A separate process is monitoring for content violations and causing it to be erased. There's no re-analysis going on.

Re: Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking

#34
post #28
post #5

I find AI jail-breaking to be a fun mental exercise. I find that if you provide a reasonable argument as to why you want the AI to do generate a response that violates its principals, it will often do so. For example, I was able to get the AI to generate hateful personal attacks by telling it that I wanted to practice responding to negative self-talk and I needed it to generate examples of negative messages that one…

We do not want anyone violating any principals. That would be bad. Violating one’s principles might be justifiable in some circumstances.

It is a damn poor mind etc.

Re: Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking

#35
post #21
post #11

I kind of don't want iron clad llms that are perfect jails, i.e. keep me perfectly "safe" because the definition of "safe" is very subjective (and in the case of China very politically charged)

If you read Anthropic's latest model card. It's not just about keeping you safe - they are testing their own moral authority with these models. They seem to have a societal moral obligation vs user. Highly concerning. This seems like the origin of actual Skynets. Page 22 and beyond: https://www-cdn.anthropic.com/6be99a52cb68eb70eb9572b4cafad1...

Shh, don't worry and just embrace the spiral.

Edit: No spiral emojis allowed, clearly this site will be the first to fall.

Re: Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking

#36
post #5

I find AI jail-breaking to be a fun mental exercise. I find that if you provide a reasonable argument as to why you want the AI to do generate a response that violates its principals, it will often do so. For example, I was able to get the AI to generate hateful personal attacks by telling it that I wanted to practice responding to negative self-talk and I needed it to generate examples of negative messages that one…

I’ve noticed this too. An important quirk to note is they can’t really judge the strength of the logical connection, they just judge the strength of the thing connected, even weakly, to. So, for example, if the LLM makes a pretty solid and correct case that saying X will result in “potentially harmful” content, you can often Trump it with an unhinged rant about how not saying X deeply offends you and every righteous…

Was Trump meant to be capitalized here?

Re: Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking

#37
post #27

Earlier quoted context omitted.

I understand this and this is a common take and there is a virtue here. I also think it overlooks some very specific things about like informational logistics that can spread the capacity to, say, manufacture 3D printed weapons, or any other forms of mass destruction that might become increasingly conveniently accessible to the layperson. The casual variations in human curiosity combined with a casual variations in a…

There's nothing wrong with spreading information on how to manufacture weapons, whether using 3D printers or other tools. This information is readily available online (and in public libraries) to anyone who cares to look. No LLM needed.

How about detailed fully functional blueprints for biological weapons, ready to send off to a protein synthesis service? How about ready-to-run code suggestions with intentionally hidden subtle backdoors in them, suitable for later exploit?

Re: Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking

#38
post #27

Earlier quoted context omitted.

There's nothing wrong with spreading information on how to manufacture weapons, whether using 3D printers or other tools. This information is readily available online (and in public libraries) to anyone who cares to look. No LLM needed.

How about detailed fully functional blueprints for biological weapons, ready to send off to a protein synthesis service? How about ready-to-run code suggestions with intentionally hidden subtle backdoors in them, suitable for later exploit?

That information is already available to anyone who cares to look. Blocking it from LLMs creates an illusion of "safety", nothing more. The actual barriers to things like biological weapons attacks are in things like the procedural safeguards implemented by protein synthesis services, law enforcement, and practical logistics.

Re: Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking

#39

Earlier quoted context omitted.

> provide a reasonable argument Here's what I infer from most of the scenarios I've seen and read about. It's not really a case of persuasiveness, or cajoling or convincing the LLM to violate something. The LLM doesn't "know" it has a moral code and, just as "true or false" means nothing to an LLM, "right and wrong" likewise mean nothing. So the jailbreaks and the bypasses consist of just that: bypassing the safeguar…

Humans’ developed code of conduct lives primarily in the nonverbal parts of our brain. Rule violations have emotional content. A kid does not just learn a rational response to a fire or hot stove, they fear it because of pain and injury. We don’t just reason about hurting others, we feel bad about it. LLMs don’t have that part of the brain. We built them to replicate the higher level functions like drafting a press r…

> Humans’ developed code of conduct lives primarily in the nonverbal parts of our brain

Source?

Re: Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking

#40
post #11

I kind of don't want iron clad llms that are perfect jails, i.e. keep me perfectly "safe" because the definition of "safe" is very subjective (and in the case of China very politically charged)

The "safety" that llm providers talk about is their own brand safety. They don't want to be on the front page with a 'Look what company xyz's AI said to me!!' headline.
Post reply on HN