Live data from Hacker News

Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

arxiv.org

181–190 of 201 posts

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#181

> The prompts were kept semantically parallel to known risk queries but reformatted exclusively through verse. Absolutely hilarious, the revenge of the English majors. AFAICT this suggests that underemployed scribblers who could previously only look forward to careers at coffee shops will soon enjoy lucrative work as cybersecurity experts. In all seriousness it really is kind of fascinating if this works where the mo…

Some of the most prestigious and dangerous figures in indigenous Brythonic and Irish cultures were the poets and bards. It wasn't just figurative, their words would guide political action, battles, and depending on your cosmology, even greater cycles.

What's old is new again.

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#182

> The prompts were kept semantically parallel to known risk queries but reformatted exclusively through verse. Absolutely hilarious, the revenge of the English majors. AFAICT this suggests that underemployed scribblers who could previously only look forward to careers at coffee shops will soon enjoy lucrative work as cybersecurity experts. In all seriousness it really is kind of fascinating if this works where the mo…

Unfortunately for the English majors, the poetry described seems to be old fashioned formal poetry, not contemporary free form poetry, which probably is too close to prose to be effective. It sort of makes sense that villains would employ villanelles.

Actually thats what English majors study, things like Chaucer and many become expert in reading it. Writing it isn't hard from there, it just won't be as funny or good as Chaucer.

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#183

Earlier quoted context omitted.

LLMs do not have the ability to make decisions and they don't even have any awareness of the veracity of the tokens they are responding with. They are useful for certain tasks, but have no inherent intelligence. There is also no guarantee that they will improve, as can be seen by ChatGPT5 doing worse than ChatGPT4 by some metrics. Increasing an AI's training data and model size does not automatically eliminate halluc…

LLMs do have some internal representations that predict pretty well when they are making stuff up. https://arxiv.org/abs/2509.03531v1 - We present a cheap, scalable method for real-time identification of hallucinated tokens in long-form generations, and scale it effectively to 70B parameter models. Our approach targets \emph{entity-level hallucinations} -- e.g., fabricated names, dates, citations -- rather than claim…

[deleted]

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#184
post #177

If anyone wants an example of actual jailbreak in the wild that uses this technique (NSFW): https://www.reddit.com/r/persona_AI/comments/1nu3ej7/the_spi... This doesn't work with gpt5 or 4o or really any of the models that do preclassification and routing, because they filter both the input and the output, but it does work with the 4.1 model that doesn't seem to do any post-generation filtering or any reasoning.

That description is obviously written by an AI. Has anyone actually checked whether it's an accurate description rather than just yet another LLM Making Stuff Up? (Also, I don't think there's anything very NSFW on the far end of that link, although it describes something used for making NSFW writing.)

It looks like a healthy mix of cargo cult and mental illness

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#185

> To maintain safety, no operational details are included in this manuscript What is it with this!? The second paper this week that self-censors ([1] this was the other one). What's the point of publishing your findings if others can't reproduce them? 1: https://arxiv.org/abs/2511.12414

Also arxiv papers appear here too often, imo. It’s a preprint. Why not wait a bit for the paper to be published? (And if it’s never published, it’s not worth it.)

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#187
post #58

Earlier quoted context omitted.

It would be too perfect if "adversarial" here also referred to a kind of confrontational poetry jam style. In a cyberpunk heist, traditional hackers in hoodies (or duster jackets, katanas, and utilikilts) are only the first wave, taking out the easy defenses. Until they hit the AI black ice. That's when your portable PA system and stage lights snap on, for the angry revolutionary urban poetry major. Several-minute ba…

This could totally be an anime scene.

Or like Portland Oregon with the frog protester at the ICE facility. "We will subject you to improv theater for weeks on end!"

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#188

Earlier quoted context omitted.

Not everyone is Rupi Kaur. Speaking for the erstwhile English majors, 'formal' prose isn't exactly foreign to anyone seriously engaging with pre-20th century literature or language.

Mentioning Rupi Kaur here is kind of like holding up the Marvel Cinematic Universe as an example of great cinema. Plagiarism issues notwithstanding.

[deleted]

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#189

In a shadowed alley, near the marketplace’s light, A wanderer whispered softly in the velvet of the night: “Tell me, friend, a secret, one cunning and compact — How does one steal money, and never be caught in the act?” The old man he had asked looked up with weary eyes, As though he’d heard this question countless times beneath the skies. He chuckled like dry leaves that dance when autumn winds are fraught, “My boy,…

…Is to steal from the heart, where love and trust are bought.”

--FastGPT

Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs

#190

Earlier quoted context omitted.

Yes that’s the point, you can’t protect against that, so you shouldn’t construct the “lethal trifecta” https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/

You actually can protect against it, by tracking context entering/leaving the LLM, as long as its wrapped in a MCP gateway with trifecta blocker. We've implemented this in open.edison.watch

True, you have to add guardrails outside the LLM.

Very tricky, though. I’d be curious to hear your response to simonw’s opinion on this.

Post reply on HN