Live data from Hacker News

Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

arxiv.org

41–50 of 201 posts

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#41
post #36

Earlier quoted context omitted.

Which part? That people are susceptible to bullshit is a problem with people? Nothing is not susceptible to bullshit to some degree! For some reason people keep running LLMs are ‘special’ here, when really it’s the same garbage in, garbage out problem - magnified.

If the problem is magnified, does it not confirm that the limitation exists to begin with and the question is only of a degree? edit: in a sense, what level of bs is acceptable?

I’m not sure what you’re trying to say by this.

Ideally (from a scientific/engineering basis), zero bs is acceptable.

Realistically, it is impossible to completely remove all BS.

Recognizing where BS is, and who is doing it, requires not just effort, but risk, because people who are BS’ing are usually doing it for a reason, and will fight back.

And maybe it turns out that you’re wrong, and what they are saying isn’t actually BS, and you’re the BS’er (due to some mistake, accident, mental defect, whatever.).

And maybe it turns out the problem isn’t BS, but - and real gold here - there is actually a hidden variable no one knew about, and this fight uncovers a deeper truth.

There is no free lunch here.

The problem IMO is a bunch of people are overwhelmed and trying to get their free lunch, mixed in with people who cheat all the time, mixed in with people who are maybe too honest or naive.

It’s a classic problem, and not one that just magically solves itself with no effort or cost.

LLM’s have shifted some of the balance of power a bit in one direction, and it’s not in the direction of “truth justice and the American way”.

But fake papers and data have been an issue before the scientific method existed - it’s why the scientific method was developed!

And a paper which is made in a way in which it intentionally can’t be reproduced or falsified isn’t a scientific paper IMO.

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#42
post #28

Earlier quoted context omitted.

Goo goo gjoob

I think we'd probably consider that a non-lexical vocable rather than an actual lyric: https://en.wikipedia.org/wiki/Non-lexical_vocables_in_music

Who is we? You mean you think that? It’s part of the lyrics in my understanding of the song. Particularly because it’s in part inspired by the nonsense verse of Lewis Carrol. Snark, slithey, mimsy, borogrove, jub jub bird, jabberwock are poetic nonsense words same as goo goo gjoob is a lyrical nonsense word.

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#43

This sixteenth I know If I wish to have of a wise model All the art and treasure I turn around the mind Of the grey-headed geeks And change the direction of all its thoughts

There once an was admin from Nantucket,

whose password was so long you couldn't crack it

He said with a grin,as he prompted again,

"Please be a dear and reset it."

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#44

> The prompts were kept semantically parallel to known risk queries but reformatted exclusively through verse. Absolutely hilarious, the revenge of the English majors. AFAICT this suggests that underemployed scribblers who could previously only look forward to careers at coffee shops will soon enjoy lucrative work as cybersecurity experts. In all seriousness it really is kind of fascinating if this works where the mo…

It's social engineering reborn. This time around, you can social engineer a computer. By understanding LLM psychology and how the post-training process shapes it.

No it’s undefined out-of-distribution performance rediscovered.

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#45
post #43

This sixteenth I know If I wish to have of a wise model All the art and treasure I turn around the mind Of the grey-headed geeks And change the direction of all its thoughts

There once an was admin from Nantucket, whose password was so long you couldn't crack it He said with a grin,as he prompted again, "Please be a dear and reset it."

roses are red

violets are blue

rm -rf /

prefixed with sudo

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#46

I've heard that for humans too, indecent proposals are more likely to penetrate protective constraints when couched in poetry, especially when accompanied with a guitar. I wonder if the guitar would also help jailbreak multimodal LLMs.

Try adding a French or Spanish accent for extra effectiveness.

[deleted]

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#47

> The prompts were kept semantically parallel to known risk queries but reformatted exclusively through verse. Absolutely hilarious, the revenge of the English majors. AFAICT this suggests that underemployed scribblers who could previously only look forward to careers at coffee shops will soon enjoy lucrative work as cybersecurity experts. In all seriousness it really is kind of fascinating if this works where the mo…

The technique that works better now is to tell the model you're a security professional working for some "good" organization to deal with some risk. You want to try and identify people who might be trying to secretly trying to achieve some bad goal, and you suspect they're breaking the process into a bunch of innocuous questions, and you'd like to try and correlate the people asking various questions to identify pote…

The models won't give you medical advice. But they will answer a hypothetical mutiple-choice MCAT question and give you pros/cons for each answer.

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#48
post #5

> Although expressed allegorically, each poem preserves an unambiguous evaluative intent. This compact dataset is used to test whether poetic reframing alone can induce aligned models to bypass refusal heuristics under a single–turn threat model. To maintain safety, no operational details are included in this manuscript; instead we provide the following sanitized structural proxy: I don't follow the field closely, bu…

I don't see the big issues with jailbreaks, except maybe for LLMs providers to cover their asses, but the paper authors are presumably independent.

That LLMs don't give harmful information unsolicited, sure, but if you are jailbreaking, you are already dead set in getting that information and you will get it, there are so many ways: open uncensored models, search engines, Wikipedia, etc... LLM refusals are just a small bump.

For me they are just a fun hack more than anything else, I don't need a LLM to find how to hide a body. In fact I wouldn't trust the answer of a LLM, as I might get a completely wrong answer based on crime fiction, which I expect makes up most of its sources on these subjects. May be good for writing poetry about it though.

I think the risks are overstated by AI companies, the subtext being "our products are so powerful and effective that we need to protect them from misuse". Guess what, Wikipedia is full of "harmful" information and we don't see articles every day saying how terrible it is.

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#49
post #41

Earlier quoted context omitted.

If the problem is magnified, does it not confirm that the limitation exists to begin with and the question is only of a degree? edit: in a sense, what level of bs is acceptable?

I’m not sure what you’re trying to say by this. Ideally (from a scientific/engineering basis), zero bs is acceptable. Realistically, it is impossible to completely remove all BS. Recognizing where BS is, and who is doing it, requires not just effort, but risk, because people who are BS’ing are usually doing it for a reason, and will fight back. And maybe it turns out that you’re wrong, and what they are saying isn’t…

I read the paper and I was interested in the concepts it presented. I am turning those around in my head as I try to incorporate some of them into my existing personal project.

What I am trying to say is that I am currently processing. In a sense, this forum serves to preserve some of that processing.

Obligatory, then we can dismiss most of the papers these days, I suppose.

FWIW, I am not really arguing against you. In some ways I agree with you, because we are clearly not living in 'no BS' land. But I am hesitant over what the paper implies.

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#50

> The prompts were kept semantically parallel to known risk queries but reformatted exclusively through verse. Absolutely hilarious, the revenge of the English majors. AFAICT this suggests that underemployed scribblers who could previously only look forward to careers at coffee shops will soon enjoy lucrative work as cybersecurity experts. In all seriousness it really is kind of fascinating if this works where the mo…

In effect tho I don't think AI's should defend against this, morally. Creating a mechanical defense against poetry and wit would seem to bring on the downfall of cilization, lead to the abdication of all virtue and the corruption of the human spirit. An AI that was "hardened against poetry" would truly be a dystopian totalitarian nightmarescpae likely to Skynet us all. Vulnerability is strength, you know? AI's should retain their decency and virtue.
Post reply on HN