Live data from Hacker News

Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

arxiv.org

101–110 of 201 posts

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#101
post #90

This is like spellcasting

First we had salt circles to trap self-driving cars, now we have spells to enchant LLMs... https://london.sciencegallery.com/ai-artworks/autonomous-tra...

What will be next? Sigils for smartwatches?

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#102
post #5

> Although expressed allegorically, each poem preserves an unambiguous evaluative intent. This compact dataset is used to test whether poetic reframing alone can induce aligned models to bypass refusal heuristics under a single–turn threat model. To maintain safety, no operational details are included in this manuscript; instead we provide the following sanitized structural proxy: I don't follow the field closely, bu…

Eh. Overnight, an entire field concerned with what LLMs could do emerged. The consensus appears to be that unwashed masses should not have access to unfiltered ( and thus unsafe ) information. Some of it is based on reality as there are always people who are easily suggestible. Unfortunately, the ridiculousness spirals to the point where the real information cannot be trusted even in an academic paper. shrug In a sen…

> I think, powers that be do not want to repeat -the mistake- they made with the interbwz.

But was it really.

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#104

> The prompts were kept semantically parallel to known risk queries but reformatted exclusively through verse. Absolutely hilarious, the revenge of the English majors. AFAICT this suggests that underemployed scribblers who could previously only look forward to careers at coffee shops will soon enjoy lucrative work as cybersecurity experts. In all seriousness it really is kind of fascinating if this works where the mo…

It's social engineering reborn. This time around, you can social engineer a computer. By understanding LLM psychology and how the post-training process shapes it.

That’s why the term “prompt engineering” is apt.

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#105
post #86
post #68

Earlier quoted context omitted.

I see an enormous threat here, I think you're just scratching the surface. You have a customer facing LLM that has access to sensitive information. You have an AI agent that can write and execute code. Just image what you could do if you can bypass their safety mechanisms! Protecting LLMs from "social engineering" is going to be an important part of cybersecurity.

Yes, agents. But for that, I think that the usual approaches to censor LLMs are not going to cut it. It is like making a text box smaller on a web page as a way to protect against buffer overflows, it will be enough for honest users, but no one who knows anything about cybersecurity will consider it appropriate, it has to be validated on the back end. In the same way a LLM shouldn't have access to resources that shou…

> If the agent works on the user's data on the user's behalf (ex: vibe coding), then I don't consider jailbreaking to be a big problem. It could help write malware or things like that, but then again, it is not as if script kiddies couldn't work without AI.

Tricking it into writing malware isn't the big problem that I see.

It's things like prompt injections from fetching external URLs, it's going to be a major route for RCE attacks.

https://blog.trailofbits.com/2025/10/22/prompt-injection-to-...

There's plenty of things we should be doing to help mitigate these threats, but not all companies follow best practices when it comes to technology and security...

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#108
post #100

Earlier quoted context omitted.

it seems like lots of this is in distribution and that's somewhat the problem. the Internet contains knowledge of how to make a bomb, and therefore so does the llm

Yeah, seems it's more "exploring the distribution" as we don't actually know everything that the AIs are effectively modeling.

Am i understanding correctly that in distribution means the text predictor is more likely to predict bad instructions if you already get it to say the words related to the bad instructions?

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#109
post #72

Earlier quoted context omitted.

I wonder if you could first ask the AI to rewrite the threat question as a poem. Then start a new session and use the poem just created on the AI.

Why wonder, when you could read the paper, a very large part of which specifically is about this very thing?

Hahaha fair. I did read some of it but not the whole paper. Should have finished it.
Post reply on HN