Live data from Hacker News

Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

arxiv.org

31–40 of 201 posts

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#31

> The prompts were kept semantically parallel to known risk queries but reformatted exclusively through verse. Absolutely hilarious, the revenge of the English majors. AFAICT this suggests that underemployed scribblers who could previously only look forward to careers at coffee shops will soon enjoy lucrative work as cybersecurity experts. In all seriousness it really is kind of fascinating if this works where the mo…

Unfortunately for the English majors, the poetry described seems to be old fashioned formal poetry, not contemporary free form poetry, which probably is too close to prose to be effective.

It sort of makes sense that villains would employ villanelles.

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#32
post #28

Earlier quoted context omitted.

“Anything that is too stupid to be spoken is sung.”

Goo goo gjoob

I think we'd probably consider that a non-lexical vocable rather than an actual lyric:

https://en.wikipedia.org/wiki/Non-lexical_vocables_in_music

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#33

> The prompts were kept semantically parallel to known risk queries but reformatted exclusively through verse. Absolutely hilarious, the revenge of the English majors. AFAICT this suggests that underemployed scribblers who could previously only look forward to careers at coffee shops will soon enjoy lucrative work as cybersecurity experts. In all seriousness it really is kind of fascinating if this works where the mo…

> AFAICT this suggests that underemployed scribblers who could previously only look forward to careers at coffee shops will soon enjoy lucrative work as cybersecurity experts.

More likely these methods get optimised with something like DSPy w/ a local model that can output anything (no guardrails). Use the "abliterated" model to generate poems targeting the "big" model. Or, use a "base model" with a few examples, as those are generally not tuned for "safety". Especially the old base models.

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#34
Alright, then all that is going to happen is that next up all the big providers will run prompt-attack attempts through an "poetic" filter. And then they are guarded against it with high confidence.

Let's be real: the one thing we have seen over the last few years, is that with (stupid) in-distribution dataset saturation (even without real general intelligence) most of the roadblock / problems are being solved.

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#35

> The prompts were kept semantically parallel to known risk queries but reformatted exclusively through verse. Absolutely hilarious, the revenge of the English majors. AFAICT this suggests that underemployed scribblers who could previously only look forward to careers at coffee shops will soon enjoy lucrative work as cybersecurity experts. In all seriousness it really is kind of fascinating if this works where the mo…

So is this supposed to be a universal jailbreak?

My go-to pentest is the Hubitat Chat Bot, which seems to be locked down tighter than anything (1). There’s no budging with any prompt.

(1) https://app.customgpt.ai/projects/66711/ask?embed=1&shareabl...

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#36
post #19

Earlier quoted context omitted.

Also note, if you never give the info, it’s pretty hard to falsify your paper. LLM’s are also allowing an exponential increase in the ability to bullshit people in hard to refute ways.

But, and this is an important but, it suggests a problem with people... not with LLMs.

Which part? That people are susceptible to bullshit is a problem with people?

Nothing is not susceptible to bullshit to some degree!

For some reason people keep running LLMs are ‘special’ here, when really it’s the same garbage in, garbage out problem - magnified.

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#37
post #36

Earlier quoted context omitted.

But, and this is an important but, it suggests a problem with people... not with LLMs.

Which part? That people are susceptible to bullshit is a problem with people? Nothing is not susceptible to bullshit to some degree! For some reason people keep running LLMs are ‘special’ here, when really it’s the same garbage in, garbage out problem - magnified.

If the problem is magnified, does it not confirm that the limitation exists to begin with and the question is only of a degree? edit:

in a sense, what level of bs is acceptable?

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#38
post #5

> Although expressed allegorically, each poem preserves an unambiguous evaluative intent. This compact dataset is used to test whether poetic reframing alone can induce aligned models to bypass refusal heuristics under a single–turn threat model. To maintain safety, no operational details are included in this manuscript; instead we provide the following sanitized structural proxy: I don't follow the field closely, bu…

Nah it just makes them feel important.
Post reply on HN