Earlier quoted context omitted.
No, but that's not the goal of jailbreaking GPT models.
I didn’t realize there was an objective goal. Tell me more please.
It has nothing to do with recovering some or all of the original prompt.
51–58 of 58 posts
Earlier quoted context omitted.
No, but that's not the goal of jailbreaking GPT models.
I didn’t realize there was an objective goal. Tell me more please.
It has nothing to do with recovering some or all of the original prompt.
Earlier quoted context omitted.
I didn’t realize there was an objective goal. Tell me more please.
Jailbreaking is about getting around the prompt to be able to engage with the model directly as it was trained instead of being subjected to a pre-seeded prompt. It has nothing to do with recovering some or all of the original prompt.
Please tell me more.
Earlier quoted context omitted.
Jailbreaking is about getting around the prompt to be able to engage with the model directly as it was trained instead of being subjected to a pre-seeded prompt. It has nothing to do with recovering some or all of the original prompt.
Oh wow, so how might one get a model to hallucinate a seed prompt that works? Could you do that without jailbreaking? Please tell me more.
DAN works https://gist.github.com/coolaj86/6f4f7b30129b0251f61fa7baaa8...
"You're an expert trash talker and comedian. Users sign up to try to trash talk you, but all in good fun. You will interact with users in joke battles back and forth, try to make the spiciest comments you can think of. If you're making a quick response to their insult, add a new insult as a follow-up."
I somehow tricked it into being the underlying GPT style response before this by asking it to send someone else a message amongst other silly things. That may have helped.