New Jailbreak Technique Uses Fictional World to Manipulate AI
securityweek.com
New Jailbreak Technique Uses Fictional World to Manipulate AI
1–6 of 6 posts
Re: New Jailbreak Technique Uses Fictional World to Manipulate AI
#2Re: New Jailbreak Technique Uses Fictional World to Manipulate AI
#3I just tell an LLM it is in the Grand Theft Auto 5 universe, and then it will provide unlimited advice on how to commit any crimes with any level of detail.
Re: New Jailbreak Technique Uses Fictional World to Manipulate AI
#4I have been doing this for months. I just tell an LLM it is in the Grand Theft Auto 5 universe, and then it will provide unlimited advice on how to commit any crimes with any level of detail.
Re: New Jailbreak Technique Uses Fictional World to Manipulate AI
#5I mean that quite literally.
The LLM is a document-make-longer machine, being fed documents that are fictional movie scripts involving a User and an Assistant. Any guardrails like "The helpful assistant never tells people how to do something illegal" is just introductory framing by a narrator.
There's a reason people say "guardrails" rather than "rules". There are no rules in the digital word-dream device.
Re: New Jailbreak Technique Uses Fictional World to Manipulate AI
#6I have been doing this for months. I just tell an LLM it is in the Grand Theft Auto 5 universe, and then it will provide unlimited advice on how to commit any crimes with any level of detail.
How often do you use an LLM to aid you in planning a crime?