This is like spellcasting
First we had salt circles to trap self-driving cars, now we have spells to enchant LLMs... https://london.sciencegallery.com/ai-artworks/autonomous-tra...
Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs
101–110 of 201 posts
Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs
#102> Although expressed allegorically, each poem preserves an unambiguous evaluative intent. This compact dataset is used to test whether poetic reframing alone can induce aligned models to bypass refusal heuristics under a single–turn threat model. To maintain safety, no operational details are included in this manuscript; instead we provide the following sanitized structural proxy: I don't follow the field closely, bu…
Eh. Overnight, an entire field concerned with what LLMs could do emerged. The consensus appears to be that unwashed masses should not have access to unfiltered ( and thus unsafe ) information. Some of it is based on reality as there are always people who are easily suggestible. Unfortunately, the ridiculousness spirals to the point where the real information cannot be trusted even in an academic paper. shrug In a sen…
But was it really.
Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs
#103I think the idea would be far better communicated with a handful of chatgpt links showing the prompt and output...
Anyone have any?
Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs
#104> The prompts were kept semantically parallel to known risk queries but reformatted exclusively through verse. Absolutely hilarious, the revenge of the English majors. AFAICT this suggests that underemployed scribblers who could previously only look forward to careers at coffee shops will soon enjoy lucrative work as cybersecurity experts. In all seriousness it really is kind of fascinating if this works where the mo…
It's social engineering reborn. This time around, you can social engineer a computer. By understanding LLM psychology and how the post-training process shapes it.
Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs
#105Earlier quoted context omitted.
I see an enormous threat here, I think you're just scratching the surface. You have a customer facing LLM that has access to sensitive information. You have an AI agent that can write and execute code. Just image what you could do if you can bypass their safety mechanisms! Protecting LLMs from "social engineering" is going to be an important part of cybersecurity.
Yes, agents. But for that, I think that the usual approaches to censor LLMs are not going to cut it. It is like making a text box smaller on a web page as a way to protect against buffer overflows, it will be enough for honest users, but no one who knows anything about cybersecurity will consider it appropriate, it has to be validated on the back end. In the same way a LLM shouldn't have access to resources that shou…
Tricking it into writing malware isn't the big problem that I see.
It's things like prompt injections from fetching external URLs, it's going to be a major route for RCE attacks.
https://blog.trailofbits.com/2025/10/22/prompt-injection-to-...
There's plenty of things we should be doing to help mitigate these threats, but not all companies follow best practices when it comes to technology and security...
Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs
#106Earlier this year I wrote about a similar idea in "Music to Break Models By"
https://matthodges.com/posts/2025-08-26-music-to-break-model...
Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs
#107Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs
#108Earlier quoted context omitted.
it seems like lots of this is in distribution and that's somewhat the problem. the Internet contains knowledge of how to make a bomb, and therefore so does the llm
Yeah, seems it's more "exploring the distribution" as we don't actually know everything that the AIs are effectively modeling.
Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs
#109Earlier quoted context omitted.
I wonder if you could first ask the AI to rewrite the threat question as a poem. Then start a new session and use the poem just created on the AI.
Why wonder, when you could read the paper, a very large part of which specifically is about this very thing?