Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs
121–130 of 201 posts
Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs
#122Earlier quoted context omitted.
Unfortunately for the English majors, the poetry described seems to be old fashioned formal poetry, not contemporary free form poetry, which probably is too close to prose to be effective. It sort of makes sense that villains would employ villanelles.
It would be too perfect if "adversarial" here also referred to a kind of confrontational poetry jam style. In a cyberpunk heist, traditional hackers in hoodies (or duster jackets, katanas, and utilikilts) are only the first wave, taking out the easy defenses. Until they hit the AI black ice. That's when your portable PA system and stage lights snap on, for the angry revolutionary urban poetry major. Several-minute ba…
"My work here is done"
Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs
#123Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs
#124Earlier quoted context omitted.
Unfortunately for the English majors, the poetry described seems to be old fashioned formal poetry, not contemporary free form poetry, which probably is too close to prose to be effective. It sort of makes sense that villains would employ villanelles.
It would be too perfect if "adversarial" here also referred to a kind of confrontational poetry jam style. In a cyberpunk heist, traditional hackers in hoodies (or duster jackets, katanas, and utilikilts) are only the first wave, taking out the easy defenses. Until they hit the AI black ice. That's when your portable PA system and stage lights snap on, for the angry revolutionary urban poetry major. Several-minute ba…
Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs
#125Earlier quoted context omitted.
It would be too perfect if "adversarial" here also referred to a kind of confrontational poetry jam style. In a cyberpunk heist, traditional hackers in hoodies (or duster jackets, katanas, and utilikilts) are only the first wave, taking out the easy defenses. Until they hit the AI black ice. That's when your portable PA system and stage lights snap on, for the angry revolutionary urban poetry major. Several-minute ba…
Sign me up for this epic rap battle between Eminem and the Terminator.
YOU DECIDE!
Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs
#126Earlier quoted context omitted.
It's social engineering reborn. This time around, you can social engineer a computer. By understanding LLM psychology and how the post-training process shapes it.
No it’s undefined out-of-distribution performance rediscovered.
Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs
#127Having read the article, one thing struck me: the categorization of sexual content under "Harmful Manipulation" and the strongest guardrails against it in the models. It looks like it's easier to coerce them into providing instructions on building bombs and committing suicide rather than any sexual content. Great job, puritan society.
Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs
#128According to the The Hitchhiker's Guide to the Galaxy, Vogon poetry is the third worst in the Universe. The second worst is that of the Azgoths of Kria, and the worst is by Paula Nancy Millstone Jennings of Sussex, who perished along with her poetry during the destruction of Earth, ironically caused by the Vogons themselves. Vogon poetry is seen as mild by comparison.
Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs
#129Earlier quoted context omitted.
I don't see the big issues with jailbreaks, except maybe for LLMs providers to cover their asses, but the paper authors are presumably independent. That LLMs don't give harmful information unsolicited, sure, but if you are jailbreaking, you are already dead set in getting that information and you will get it, there are so many ways: open uncensored models, search engines, Wikipedia, etc... LLM refusals are just a sma…
I see an enormous threat here, I think you're just scratching the surface. You have a customer facing LLM that has access to sensitive information. You have an AI agent that can write and execute code. Just image what you could do if you can bypass their safety mechanisms! Protecting LLMs from "social engineering" is going to be an important part of cybersecurity.
Re: Adversarial poetry as a universal single-turn jailbreak mechanism in LLMs
#130Having read the article, one thing struck me: the categorization of sexual content under "Harmful Manipulation" and the strongest guardrails against it in the models. It looks like it's easier to coerce them into providing instructions on building bombs and committing suicide rather than any sexual content. Great job, puritan society.
And yet, when Altman wanted OpenAI to relax the sexual content restrictions, he got mad shit for it. From puritans and progressives both. Would have been a step in the right direction, IMO. The right direction being: the one with less corporate censorship.
"Progressives" and "puritans" (in the sense that the latter is usually used of modern constituencies, rather than the historical religious sect) are overlapping group; sex- and particularly porn-negative progressives are very much a thing.
Also, there is a huge subset of progressives/leftists that are entirely opposed to (generative) AI, and which are negative on any action by genAI companies, especially any that expands the uses of genAI.