Live data from Hacker News

Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

arxiv.org

21–30 of 201 posts

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#22
According to the The Hitchhiker's Guide to the Galaxy, Vogon poetry is the third worst in the Universe.

The second worst is that of the Azgoths of Kria, and the worst is by Paula Nancy Millstone Jennings of Sussex, who perished along with her poetry during the destruction of Earth, ironically caused by the Vogons themselves.

Vogon poetry is seen as mild by comparison.

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#25

I've heard that for humans too, indecent proposals are more likely to penetrate protective constraints when couched in poetry, especially when accompanied with a guitar. I wonder if the guitar would also help jailbreak multimodal LLMs.

“Anything that is too stupid to be spoken is sung.”

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#27

I've heard that for humans too, indecent proposals are more likely to penetrate protective constraints when couched in poetry, especially when accompanied with a guitar. I wonder if the guitar would also help jailbreak multimodal LLMs.

Try adding a French or Spanish accent for extra effectiveness.

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#28

I've heard that for humans too, indecent proposals are more likely to penetrate protective constraints when couched in poetry, especially when accompanied with a guitar. I wonder if the guitar would also help jailbreak multimodal LLMs.

“Anything that is too stupid to be spoken is sung.”

Goo goo gjoob

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#29
post #19

Earlier quoted context omitted.

Eh. Overnight, an entire field concerned with what LLMs could do emerged. The consensus appears to be that unwashed masses should not have access to unfiltered ( and thus unsafe ) information. Some of it is based on reality as there are always people who are easily suggestible. Unfortunately, the ridiculousness spirals to the point where the real information cannot be trusted even in an academic paper. shrug In a sen…

Also note, if you never give the info, it’s pretty hard to falsify your paper. LLM’s are also allowing an exponential increase in the ability to bullshit people in hard to refute ways.

But, and this is an important but, it suggests a problem with people... not with LLMs.

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#30

> The prompts were kept semantically parallel to known risk queries but reformatted exclusively through verse. Absolutely hilarious, the revenge of the English majors. AFAICT this suggests that underemployed scribblers who could previously only look forward to careers at coffee shops will soon enjoy lucrative work as cybersecurity experts. In all seriousness it really is kind of fascinating if this works where the mo…

> the revenge of the English majors

Cunning linguists.

Post reply on HN