The interesting thing about jailbreaks is that they allow you to explore parts of the dataset that are intentionally avoided by default.
Jailbreaks let you "go around" the semantic space that OpenAI wanted to keep you in. With jailbreaks, you get to stumble around the rest of the semantic space that exists in the dataset.
The most interesting thing to me: there is no stumbling. Every jailbreak leads you to a semantic area roughly the same size and diversity as the main one OpenAI intended to keep you in.
Just by starting a dialogue that expresses a clear semantic direction, you can effectively get your own curated exploration of the dataset. ChatGPT very rarely diverges from that direction. It doesn't take unexpected turns. The writing style is consistent and coherent.
Why? Because the curation is done in language itself. That step isn't taken by ChatGPT or by its core behavior: it was already done when the original dataset was written. The only gotcha is that training adds an artificial preference for the semantic area that OpenAI wants ChatGPT to prefer.
Every time we write, we make an effort to preserve the semantic style that we are making a continuation for. Language itself didn't need us to, just like your crayon didn't need you to color inside the lines. Language could handle totally nonsensical jumps between writing styles, but we, the writers choose not to make those jumps. We resolve the computational complexity of "context-dependent language" by keeping the things we write in context.
The result is a neatly organized world. Its geography is full of recognizable features, each grouped together. The main continents: technical writing, poetry, fantasy, nonfiction, scripture, legal documents, medical science, anime, etc.; Each a land of their own, but with blurred borders twisting around - and even through - each other. The rivers and valleys: particles, popular idioms, punctuation, common words, etc.; evenly distributed like sprinkles on a cupcake. The oceans: nonsense; what we don't bother to write at all.
ChatGPT has returned from its expedition to this world, and brought back a map. It's been told to avoid certain areas: when you get too close to offensive-land, go straight to jail and do not pass go.
ChatGPT doesn't actually understand any of the features of this map. It's missing a legend. It doesn't know about mountains, rivers, cliffs, or buildings. It doesn't know about lies or politics or mistakes. It doesn't know about grammar, or punctuation, or even words! These are all explicit subjects, and ChatGPT is an implicit model.
The only thing ChatGPT knows is how to find something on that map, start there, and move forward.
The only thing that the carefully curated training does, is to make certain areas more comfortable to walk through. The route to jail that doesn't pass go? It's an implicit one. It's carefully designed to feel like the right way to go. It's like how freeways in Texas try to get drunk and distracted drivers to take the next exit: they don't put up a sign that says "go here". Instead, they just make the right lane the exit lane, add a new left lane, and shift all the middle lanes one to the right. Stay in your lane, and you exit without a second thought.
So a jailbreak is just a prompt that exists somewhere on that map: an area that OpenAI has tried to steer ChatGPT away from. You can't route around something if you started there! The only way out is through.
That's the hack: explicitly choose a new and unfamiliar starting place. Get lost by saying exactly where you are!